Big data diagnosis method and system based on auscultation data
By combining multi-scale chaotic measurement and temporal topology analysis with causal relationship analysis of neural networks and Bayesian networks, the problem of pathological event capture and causal analysis in auscultatory diagnosis is solved, realizing a high-precision and interpretable auscultatory diagnostic method.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUIZHOU UNIV
- Filing Date
- 2025-12-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing auscultation diagnostic methods are unable to accurately capture complex pathological events and lack causal analysis capabilities. Furthermore, deep learning models lack interpretability and the ability to integrate medical knowledge, resulting in insufficient usability and reliability of diagnostic results.
Multi-scale chaotic measures are used to screen pathologically sensitive fragments. Features are extracted by combining temporal convolution, time-frequency convolution and temporal modeling neural network structures. Temporal topology is constructed and causal relationships are analyzed by Bayesian networks. Pathological acoustic causal intensity indicators and prior knowledge are introduced to optimize LSTM prediction.
It significantly improves the accuracy, stability, and interpretability of auscultatory diagnosis, enhances the characteristic sensitivity to complex heart and lung sounds, and improves the clinical consistency of diagnosis.
Smart Images

Figure CN121964097A_ABST
Abstract
Description
A Big Data Diagnostic Method and System Based on Auscultation Data Technical Field
[0001] This invention relates to the field of big data diagnostic technology, specifically to a big data diagnostic method and system based on auscultation data. Background Technology
[0002] Auscultation, as one of the most commonly used non-invasive examination methods in clinical practice, has long relied on physicians' subjective experience in interpreting heart or lung sounds. However, in real clinical scenarios, auscultation signals exhibit significant complexity: they not only include periodic fluctuations caused by physiological rhythms but also transient disturbances triggered by pathological events, respiratory background noise, surface friction sounds, and equipment artifacts. These non-stationary characteristics make traditional segmentation methods based on fixed windows, energy thresholds, and empirical rules difficult to accurately capture pathologically sensitive segments. Furthermore, cardiopulmonary diseases often manifest as coupled changes in multiple acoustic features, such as high-frequency disturbances caused by valvular abnormalities, persistent whistling due to bronchial stenosis, and crackling sounds induced by alveolar inflammation. Single time-domain or frequency-domain analysis methods are insufficient to comprehensively characterize these multimodal acoustic structures.
[0003] In recent years, deep learning has become increasingly popular for auscultatory analysis. However, existing methods generally rely on manual slides and fixed-scale feature extraction, which cannot handle the random occurrence of complex pathological events at different time scales. Furthermore, deep models often lack interpretability, failing to describe the causal relationships, dependencies, and disease-related structures between different acoustic events, nor can they integrate the positive and negative correlations of diseases from real-world medical knowledge. Therefore, the diagnostic results still have significant shortcomings in terms of clinical usability and reliability.
[0004] Furthermore, the diagnostic process of auscultation exhibits a clear causal chain characteristic: local acoustic events can influence subsequent rhythmic structures, and typical acoustic manifestations of different diseases also show comorbidity or mutual exclusion. However, existing studies generally only perform feature-based classification and lack the ability to quantify, analyze, and track the causal importance of each segment in a time series. Summary of the Invention
[0005] In view of the above-mentioned problems, the present invention is proposed.
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a big data diagnostic method based on auscultation data, comprising:
[0007] The original auscultation sequence of the target individual is obtained using a stethoscope. The auscultation sequence is preprocessed, and pathologically sensitive fragments in local time windows are screened based on multi-scale chaos measure to obtain fragment sequences.
[0008] The segment sequence is feature-encoded using an encoder to obtain an auscultatory feature sequence;
[0009] A temporal topology is constructed based on the auscultation feature sequence, and the pathological acoustic causal intensity index characterizing disease correlation is calculated for each auscultation feature by analyzing the causal relationship between the auscultation features.
[0010] The auscultatory feature sequence with the pathological acoustic causal intensity index is converted into a causal state sequence under each disease modality, and the disease diagnosis of the target individual is performed through the state transition between the causal state sequences.
[0011] As a preferred embodiment of the big data diagnostic method based on auscultation data described in this invention, the original auscultation sequence includes a target individual auscultation time sequence continuously collected by a stethoscope at a fixed sampling rate, reflecting the amplitude changes of heart sounds or lung sounds in the time dimension.
[0012] The preprocessing includes amplitude normalization, pulse artifact removal, steady-state background noise suppression, and temporal boundary detection of the original auscultation sequence to construct a basic processing sequence for subsequent multi-scale chaotic measurement analysis.
[0013] As a preferred embodiment of the big data diagnostic method based on auscultation data described in this invention, the multi-scale chaos measure includes downsampling the original auscultation sequence based on n preset different downsampling ratios to generate n downsampling subsequences; the set of all sequences is denoted as: Using a sliding time window in time sequence, calculate the chaotic feature value of each sliding time window in the original auscultation sequence and each of the downsampled subsequences, respectively.
[0014] in, This represents a downsampled subsequence with a downsampling factor of n;
[0015] On the timeline, according to arrive In each round, time is sliced according to the chaotic feature value in each sequence; after the previous round of slicing is completed, the next round of slicing is performed on each slice respectively.
[0016] The slicing steps for each round are as follows:
[0017] Step 1: In any slice of time period in the previous round, extract the downsampled subsequence corresponding to the current round, and form a single-scale chaotic sequence representing the local complexity change of the sequence based on the chaotic feature value of each sliding time window;
[0018] Step 2: In the single-scale chaotic sequence, the relative deviation of the chaotic feature values of adjacent windows is judged according to the time sequence: if the relative deviation exceeds the preset threshold of the corresponding downsampling rate, the end of the previous sliding window is identified as the segmentation point.
[0019] Step 3: Based on the stated dividing points, divide the slices from the previous round; after division, all slices for this round are obtained.
[0020] After completing the slicing in this round, the slicing in the next round is performed; after repeating n rounds, a fragment sequence of the original auscultation sequence is obtained.
[0021] As a preferred embodiment of the big data diagnostic method based on auscultation data described in this invention, the encoder includes inputting the time-domain signal in the slice into a neural network feature extraction structure, performing nonlinear mapping on the morphological features, frequency band energy distribution features and short-time dynamic change features of each segment to obtain a multidimensional local feature vector of the slice.
[0022] The neural network feature extraction structure includes a time-domain convolution branch that performs convolution operations on the original waveform, a time-frequency convolution branch that performs convolution operations on the time-frequency representation obtained by time-frequency transformation, and a time-series modeling branch for modeling the time dependencies within the segment.
[0023] The temporal convolution branch is used to extract the morphological features of the original waveform;
[0024] The time-frequency convolution branch is used to extract energy structure and frequency band distribution features;
[0025] The temporal modeling branch is used to capture dynamic evolution features and extract short-term dynamic change features.
[0026] As a preferred embodiment of the big data diagnostic method based on auscultation data described in this invention, the temporal topology structure includes: using slices as nodes, establishing edge relationships between any two nodes, and using a Bayesian network to generate probability distributions of the two nodes in each disease to represent the causal relationship between them.
[0027] Among them, for any two nodes, the sum of the probabilities of different diseases, health states and unidentifiable conditions is in the range of 1.
[0028] For any node in the temporal topology, after evaluating its ontology value, the corresponding slice is used as the starting point for evolution. Evolutionary analysis is performed in both the positive and negative directions of the time axis, and the overall value of the node is generated based on the correlation strength of the node during the evolutionary analysis. If the ontology value or the overall value is lower than the corresponding threshold, the slice corresponding to the node is simulated by merging with the previous slice and the next slice, respectively. The information gain of the ontology value and the overall value is calculated when merging forward and backward. The merging direction is selected with the goal of maximizing the weighted sum of the information gain, and the slice is merged according to the merging direction.
[0029] The ontology value includes, for node i, generating a value coefficient based on the edge relationship between i and any node j: ;in, The standard deviation of different probabilities in the edge relationship between i and j is used to measure the unevenness of probabilities. This represents the probability that the edge relationship between i and j cannot be identified.
[0030] We obtain the final ontology value of i by performing a weighted summation on all edge relationships: ;
[0031] Where J represents the number of nodes; The weight between node i and node j is generated by proportionally decaying the initial value A in the sequence of nodes in both the forward and backward directions to generate a reference value for each node, and then normalizing it.
[0032] The evolutionary analysis includes inputting the node sequence and edge relationships into a pre-trained LSTM; the LSTM slides through the node sequence from front to back, and at each node, based on the edge relationships between the preceding nodes, generates the probabilities of different diseases, health states, and unidentifiable conditions, thus obtaining a predicted probability distribution.
[0033] Using the labels in the predicted probability distribution as dimensions, and the probability of each label as the value in the dimension, multidimensional features are generated. The features of each dimension are fitted according to the order of the nodes. After calculating the relative deviation of node i in the fitting results of each dimension, the (1-relative deviation) of each dimension is summed to obtain the overall value of node i.
[0034] As a preferred embodiment of the big data diagnostic method based on auscultation data described in this invention, in the temporal topology structure after merging slices, the overall value of each node is summed and then divided by the number of nodes to obtain the pathological acoustic causality intensity index.
[0035] As a preferred embodiment of the big data diagnostic method based on auscultation data described in this invention, the causal state sequence includes the feature sequence of each node in each dimension of the time topology structure after slicing and merging.
[0036] The state transition includes learning the positive and negative correlations between different diseases based on prior probabilities; and analyzing whether the changes in LSTM prediction probabilities between feature sequences of different dimensions conform to the positive and negative correlations.
[0037] If a matching correlation exists, then for diseases whose LSTM prediction probability increases in the sequence, the attention weights obtained after linear transformation according to the prior probability corresponding to the correlation are used for attention enhancement in the LSTM to achieve the final output.
[0038] A big data diagnostic system based on auscultation data using any of the methods described in this invention, wherein: a data acquisition unit acquires the original auscultation sequence of the target individual using a stethoscope, preprocesses the auscultation sequence, and, based on multi-scale chaos measure, screens pathologically sensitive segments in local time windows to obtain segment sequences;
[0039] The encoding unit uses an encoder to encode the features of the segment sequence to obtain an auscultatory feature sequence;
[0040] The calculation unit constructs a temporal topology based on the auscultation feature sequence, and calculates the pathoacoustic causal intensity index representing the disease correlation in each auscultation feature by analyzing the causal relationship between the auscultation features.
[0041] The migration unit converts the auscultatory feature sequence with the pathological acoustic causal intensity index into a causal state sequence for each disease mode, and performs disease diagnosis for the target individual through state migration between the causal state sequences.
[0042] A computer device includes: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, it implements the steps of the method described in any one of the present invention.
[0043] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of the present invention.
[0044] The beneficial effects of this invention are as follows: This invention adaptively slices the original auscultation sequence using multi-scale chaotic measures, achieving precise localization of pathologically sensitive time periods, which significantly improves the pathological event capture rate compared to traditional fixed-window segmentation. Employing a three-branch neural network structure combining temporal convolution, time-frequency convolution, and temporal modeling, the morphological, energy, and dynamic features of each segment can be fully expressed, enhancing the feature sensitivity to complex heart and lung sounds. By constructing a temporal topology and introducing causal edge relationships based on Bayesian networks, the pathological influence paths between different segments can be characterized, making the diagnostic process interpretable and enabling chain-like reasoning. Simultaneously, based on dual indicators of node ontological value and overall value, this invention can automatically merge unstable segments, improving the robustness of the topology. Furthermore, by introducing prior knowledge of the positive and negative correlations of diseases, attention enhancement is applied to the LSTM prediction results, effectively improving the clinical consistency of the diagnosis. This invention significantly improves the accuracy, stability, and interpretability of auscultation diagnosis. Attached Figure Description
[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 is an overall flowchart of a big data diagnostic method based on auscultation data provided in the first embodiment of the present invention. Detailed Implementation
[0047] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0048] Referring to Figure 1, an embodiment of the present invention provides a big data diagnostic method based on auscultation data, including:
[0049] S1: Obtain the original auscultation sequence of the target individual using a stethoscope, preprocess the auscultation sequence, and screen pathologically sensitive segments in the local time window based on multi-scale chaos measure to obtain segment sequences.
[0050] Furthermore, the original auscultation sequence includes a sequence of auscultation time-series signals of the target individual continuously acquired by a stethoscope at a fixed sampling rate, reflecting the amplitude changes of heart sounds or lung sounds in the time dimension. The preprocessing includes amplitude normalization, pulse artifact removal, steady-state background noise suppression, and temporal boundary detection of the original auscultation sequence to construct a basic processing sequence for subsequent multi-scale chaotic measurement analysis. Amplitude normalization ensures that auscultation signals under different acquisition conditions have consistent dimensions, avoiding interference from differences in device gain and contact force on chaotic features; pulse artifact removal eliminates non-physiological sudden interferences such as wire friction and sensor vibration, preventing them from being misjudged as pathological mutations; steady-state background noise suppression reduces the rise of the local complexity baseline due to respiratory and environmental noise, improving the responsiveness of chaotic measures to weak pathological signals; and temporal boundary detection ensures that the analysis window is consistent with the physiological cycle of heart sounds or lung sounds, avoiding feature distortion caused by slices spanning different physiological stages. Through the above processing, a basic processing sequence with a clear structure and controlled noise can be constructed, providing a reliable data foundation for subsequent segmentation, feature extraction and causal diagnosis, and enabling more accurate identification of abnormal acoustic events.
[0051] It should be noted that the multi-scale chaos measure includes downsampling the original auscultation sequence based on n preset different downsampling ratios to generate n downsampling subsequences; the set of all sequences is denoted as: Using a sliding time window over time, the chaotic feature value of each sliding time window is calculated in both the original auscultation sequence and each downsampled subsequence. This represents a downsampled subsequence with a downsampling factor of n.
[0052] On the timeline, according to arrive In each round, time is sliced according to the chaotic feature value in each sequence; after the previous round of slicing is completed, the next round of slicing is performed on each slice.
[0053] The slicing steps for each round are as follows:
[0054] Step 1: In any slice of time period in the previous round, extract the downsampled subsequence corresponding to the current round, and form a single-scale chaotic sequence representing the local complexity change of the sequence based on the chaotic feature value of each sliding time window.
[0055] Step 2: In the single-scale chaotic sequence, the relative deviation of the chaotic feature values of adjacent windows is judged according to the time sequence: if the relative deviation exceeds the preset threshold of the corresponding downsampling rate, the end of the previous sliding window is identified as the segmentation point.
[0056] Step 3: Based on the stated dividing point, divide the slices from the previous round; after division, all slices for this round are obtained.
[0057] After completing the slicing in this round, the slicing in the next round is performed; after repeating n rounds, a fragment sequence of the original auscultation sequence is obtained.
[0058] It's important to understand that, to achieve precise localization of pathological events in auscultatory signals, this invention, within a multi-scale chaotic measurement framework, slices sequences generated at different downsampling rates layer by layer to capture the complexity mutations of heart or lung sounds at different time scales. The design aims to: high-magnification downsampling highlights low-frequency structural changes, suitable for identifying long-term, slowly changing pathological events; low-magnification or the original sequence retains more high-frequency details, aiding in the detection of transient acoustic anomalies. By pressing L... n →L1 slicing proceeds sequentially, refining fragment segmentation from coarse to fine, avoiding omissions or missegments of pathological fragments at a single time scale. In each round of slicing, the downsampled sequence at the current scale is mapped to a single-scale chaotic sequence, with the chaotic feature value of each sliding window reflecting the local complexity level. The "relative deviation" between two adjacent windows is used to determine whether there is a sudden change in complexity. It is calculated by comparing the change in the chaotic feature value of the later window relative to the earlier window, and dividing this change by the feature value of the earlier window to obtain a deviation value representing the proportion of change. When this deviation exceeds a threshold set by the corresponding downsampling rate, a significant change in complexity is considered to have occurred at the current scale, thus marking the end of the previous window as a segmentation point. Based on these segmentation points, the slices from the previous round are further refined. Through multiple iterations, a fragment sequence that simultaneously considers coarse-grained trends and fine-grained pathological details is finally obtained, providing a precise temporal structure basis for subsequent feature encoding and causal analysis.
[0059] S2: Use an encoder to perform feature encoding on the segment sequence to obtain an auscultation feature sequence.
[0060] The encoder includes inputting the time-domain signal in the slice into a neural network feature extraction structure, performing nonlinear mapping on the morphological features, frequency band energy distribution features and short-time dynamic change features of each segment, and obtaining a multidimensional local feature vector of the slice.
[0061] The neural network feature extraction structure includes a time-domain convolution branch that performs convolution operations on the original waveform, a time-frequency convolution branch that performs convolution operations on the time-frequency representation obtained by time-frequency transformation, and a time-series modeling branch for modeling the time dependencies within the segment.
[0062] The temporal convolution branch is used to extract the morphological features of the original waveform.
[0063] The time-frequency convolution branch is used to extract energy structure and frequency band distribution features.
[0064] The temporal modeling branch is used to capture dynamic evolution features and extract short-term dynamic change features.
[0065] To achieve a comprehensive representation of the internal acoustic structure of pathological fragments, a multi-branch neural network encoder is used to encode the features of the slices, solving problems such as rapid morphological changes, complex frequency band structure, and subtle local dynamic changes in the original auscultation signal. This enables subsequent temporal topology modeling and causal inference to be carried out based on structured and comparable high-dimensional features.
[0066] Specifically, the original time-domain waveform in the segment first enters the time-domain convolution branch to extract waveform variation features such as heart sound morphology, rising edge of pop sounds, and periodicity of whistling sounds. Simultaneously, the spectrum of the segment after time-frequency transformation is input into the time-frequency convolution branch to extract frequency band distribution features such as high-frequency structures of valvular abnormalities and energy bandwidth expansion caused by pulmonary airflow turbulence. Further, the segment is input into the time-series modeling branch to capture short-term dynamic features that are time-dependent, such as instantaneous pops, periodic changes in wheezing, and arrhythmias. Through multi-branch nonlinear mapping, a multi-dimensional local feature vector covering morphology, frequency band, and dynamic evolution can be obtained.
[0067] In order to enable the subsequent construction of the temporal topology to simultaneously utilize the "acoustic originality of the pathological fragments themselves" and the "structured representation extracted by the model", the encoded local feature vectors and the original temporal signals of the slices or their compressed forms are used together as the input features of the nodes. This allows the nodes to retain real acoustic information and have a high-level expression suitable for causal relationship analysis and probabilistic modeling, thereby improving the reliability of subsequent causal strength calculation, node value assessment and state transition diagnosis.
[0068] S3: Construct a temporal topology based on the auscultation feature sequence, and calculate the pathological acoustic causal intensity index that characterizes the disease correlation in each auscultation feature by analyzing the causal relationship between the auscultation features.
[0069] The temporal topology structure includes using slices as nodes, establishing edge relationships between any two nodes, and using a Bayesian network to generate probability distributions of the two nodes for each disease to represent the causal relationship between them.
[0070] For any two nodes, the sum of the probabilities of different diseases, health states, and unidentifiable conditions is in the range of 1.
[0071] For any node in the temporal topology, after evaluating its ontology value, the corresponding slice is used as the starting point for evolution. Evolutionary analysis is performed in both the positive and negative directions of the time axis, and the overall value of the node is generated based on the correlation strength of the node during the evolutionary analysis. If the ontology value or the overall value is lower than the corresponding threshold, the slice corresponding to the node is simulated by merging with the previous slice and the next slice, respectively. The information gain of the ontology value and the overall value is calculated when merging forward and backward. The merging direction is selected with the goal of maximizing the weighted sum of the information gain, and the slice is merged according to the merging direction. This process is repeated until the threshold is met.
[0072] Two thresholds are used to analyze whether the slice is effective. If it is ineffective, it is merged into other relatively inefficient information nearby to avoid contamination caused by merging into efficient slices.
[0073] To explain, to ensure that the temporal topology accurately reflects the pathological evolution in the auscultation sequence, after constructing the causal probabilities between nodes, a segment adaptive optimization mechanism based on a dual index of "ontology value" and "overall value" is introduced. This mechanism utilizes the causal relationships between nodes obtained through a Bayesian network to quantify the importance of each segment in the local causal chain, and evaluates its contribution to the global pathological path through forward and backward evolutionary analysis. This identifies low-value segments caused by excessively short segments, unstable features, or local noise interference. When the ontology value or overall value of a segment is below a threshold, it indicates a lack of sufficient support in the causal structure. To prevent these "weak segments" from directly entering subsequent disease discrimination and causing inference errors, they are simulated and merged into adjacent segments. The magnitude of the value increase after merging is used as information gain, and the side with the largest increase is selected from both forward and backward merging. This allows the segments to be reconstructed at multiple scales until the value of all segments reaches a stable threshold. By setting two different levels of thresholds, we can avoid erroneously merging noisy fragments into high-value fragments, prevent contamination of key pathological information, and ensure that the slices ultimately used for causal strength assessment and state transition analysis are all effective fragments that have undergone quality filtering, forming a time topology network with a clear structure and reliable causality.
[0074] Specifically, the ontology value includes generating a value coefficient for node i based on the edge relationship between i and any node j: ;in, The standard deviation of different probabilities in the edge relationship between i and j is used to measure the unevenness of probabilities. This represents the probability that the edge relationship between i and j cannot be identified.
[0075] We obtain the final ontology value of i by performing a weighted summation on all edge relationships: .
[0076] Where J represents the number of nodes; The weight between node i and node j is generated by proportionally decaying the initial value A in the sequence of nodes in both directions, and then normalizing it.
[0077] The evolutionary analysis includes inputting the node sequence and edge relationships into a pre-trained LSTM; the LSTM slides through the node sequence from front to back, and at each node, based on the edge relationships between the preceding nodes, generates the probabilities of different diseases, health states, and unidentifiable conditions, thus obtaining a predicted probability distribution.
[0078] The labels (different diseases, health statuses, and unidentifiable labels) in the predicted probability distribution are used as dimensions, and the probability of each label is used as the value in the dimension to generate multidimensional features. The features of each dimension are fitted according to the order of the nodes. After calculating the relative deviation of node i in the fitting results of each dimension, the (1-relative deviation) of each dimension is summed to obtain the overall value of node i.
[0079] To ensure that each node in the temporal topology accurately reflects the importance of pathological acoustics at both the local and global levels, ontological value is used to evaluate the stability and reliability of the causal probability distribution between a node and all other nodes. If the causal relationship of a node exhibits high heterogeneity (large standard deviation) or high uncertainty (high probability of being unidentifiable), its value coefficient is reduced, thus avoiding the direct use of noisy or structurally unstable segments for disease inference. Simultaneously, the overall value obtained through evolutionary analysis is used to measure the temporal consistency of a node within the entire sequence, i.e., whether the node matches the overall pathological evolution trend of the sequence. Specifically, LSTM generates probability sequences for different diseases, healthy states, and unidentifiable states based on the causal relationships of preceding nodes. Each label of these probability sequences is used as an independent dimension, and curve fitting is performed on each dimension along the node sequence. The fitting result represents the "normal trend" of the sequence under that dimension.
[0080] The "relative deviation" of a node is the proportion of its actual probability value relative to the predicted value of the fitted curve. It is calculated by taking the absolute value of the difference between the actual value and the fitted value, and normalizing it to the fitted value. The resulting value represents the relative magnitude of the deviation in that dimension. A larger relative deviation indicates that the node's behavior in that dimension is less consistent with the overall trend. The score is calculated as (1 - relative deviation), and these scores are accumulated to obtain the overall value, which accurately reflects the node's stability and contribution to the global causal chain. Through these dual-value indicators, unreliable nodes can be accurately identified and merged and reconstructed in subsequent steps, ensuring that all segments participating in the causal strength calculation are of high quality.
[0081] Furthermore, in the temporal topology structure after slicing and merging, the summation of the overall value of each node is divided by the number of nodes to obtain the pathological acoustic causality intensity index.
[0082] S4: The auscultatory feature sequence with the pathological acoustic causal intensity index is converted into a causal state sequence under each disease mode, and the disease diagnosis of the target individual is performed through the state transition between the causal state sequences.
[0083] The causal state sequence includes the feature sequence of a node in each dimension within the temporal topology after slice merging. These are sequences representing different dimensions.
[0084] The state transition includes learning the positive and negative correlations between different diseases based on prior probabilities; and analyzing whether the changes in LSTM prediction probabilities between feature sequences of different dimensions conform to the positive and negative correlations.
[0085] If a matching correlation exists, then for diseases whose LSTM prediction probability increases in the sequence, the attention weights obtained after linear transformation according to the prior probability corresponding to the correlation are used for attention enhancement in the LSTM to achieve the final output.
[0086] It's important to understand that prior probabilities represent the model's initial understanding of a disease, typically provided by historical data or expert knowledge. However, directly using prior probabilities as attention weights can lead to an over-biased approach to certain diseases. Therefore, these weights can be adjusted through linear transformations (such as multiplying by a reduction factor less than 1) to ensure their moderate influence in the LSTM model. For example, if the prior probability of disease A is 0.8 and that of disease B is 0.2, these probabilities can be reduced (e.g., multiplied by 0.5) to minimize their impact on the model.
[0087] In LSTM, the attention mechanism weights the input features at each time step based on the current prediction and historical data. By adjusting the attention weights of the LSTM using the reduced prior probabilities (as a result of linear transformation), the LSTM can be effectively made to focus on specific disease features without being overly biased towards certain diseases.
[0088] To improve the reliability and medical consistency of disease diagnosis results, this invention, after obtaining the pathological acoustic causal intensity index of the node, transforms the auscultation feature sequence into a causal state sequence under each disease modality, and introduces a state transfer mechanism based on prior knowledge: by utilizing the comorbidity, mutual exclusion and causal relationships between diseases in medical knowledge and historical data, the deep model not only depends on the sound features themselves, but also follows the disease structure that actually exists in the medical field, thereby suppressing the random bias and local misjudgment of the model.
[0089] The concept of "learning the positive and negative correlations between different diseases based on prior probability" refers to the understanding, through statistical analysis of a large number of case samples or clinical experience, that certain diseases share a common tendency to occur (positive correlation), such as asthma and airway stenosis; while other diseases exhibit mutually exclusive or competitive relationships (negative correlation), such as normal heart sounds and various pathological murmurs. By calculating the probability of the correlation between different diseases, and for each correlation, using the relative deviation rate of the correlation strength from the standard correlation strength as a reduction rate, the prior probability is obtained by multiplying the probability of the correlation with the corresponding reduction rate. This allows us to obtain the prior influence of each disease on another disease.
[0090] "Correlation strength" is used to quantify the strength of the correlation between the two disease sequences in this scheme. It is a linear measure of the co-occurrence trend (positive correlation) or mutual exclusion trend (negative correlation) between the diseases. This strength is achieved by linearly fitting the probability sequences formed by the changes in the conditional probabilities of the two diseases over time or in case samples, and using the slope of the fitted line as the measure of correlation strength. A positive slope indicates that the two diseases have a co-occurrence linear trend, a negative slope indicates that the two diseases are mutually exclusive, and a slope close to zero indicates a weak correlation or no significant trend. The "standard correlation strength" is the standard slope for this correlation, which can be obtained through preset or training.
[0091] The "relative deviation rate" measures the degree to which the correlation strength of a disease pair deviates from the standard correlation strength. It is a quantitative evaluation of whether the association level of the disease pair is abnormally strong or abnormally weak. It is calculated as follows: First, the difference between the correlation strength of the disease pair and the standard correlation strength is calculated to represent the absolute deviation. Then, this absolute deviation is divided by the standard correlation strength to obtain the deviation ratio relative to the baseline level, which is the "relative deviation rate." A large relative deviation rate indicates that the linear correlation of the disease pair significantly deviates from the typical correlation level of the disease set, regardless of whether the deviation is positive or negative. A small relative deviation rate indicates that the correlation of the disease pair is consistent with the correlation of the overall disease population and is not an abnormal association. By using the relative deviation rate as a reduction rate, the original prior probability can be linearly decayed, preventing disease associations deviating from the typical level from causing excessive influence during state transitions, thereby maintaining the stability and rationality of the model when prior knowledge is introduced.
[0092] The term "existence of a consistent correlation" means that when the direction of change in the predicted probability of LSTM is consistent with the positive or negative trend recorded in the prior, for example, if the predicted probability of disease A increases and A is positively correlated with B in the prior, then the relationship is valid; or if the probability of A increases while B is negatively correlated with A, and the predicted probability of B decreases, then it also belongs to the category of consistent trend.
[0093] When the trend holds true, prior probabilities are introduced as attention weights for diseases with increasing predicted probabilities in the LSTM. However, to avoid over-biasing towards high-probability diseases, the prior probabilities are linearly reduced according to a decay rate to keep their influence moderate. The reduced prior weights participate in the LSTM's attention mechanism, giving extra attention to specific disease modalities. This allows the model to fully utilize medical priors during state transitions between diseases and improves the consistency and accuracy of the final diagnosis.
[0094] On the other hand, this embodiment also provides a big data diagnostic system based on auscultation data, which includes:
[0095] The acquisition unit uses a stethoscope to acquire the original auscultation sequence of the target individual, preprocesses the auscultation sequence, and uses multi-scale chaos measure to screen pathologically sensitive segments in local time windows to obtain segment sequences.
[0096] The encoding unit uses an encoder to perform feature encoding on the segment sequence to obtain an auscultatory feature sequence.
[0097] The calculation unit constructs a temporal topology based on the auscultation feature sequence, and calculates the pathological acoustic causal intensity index representing the disease correlation in each auscultation feature by analyzing the causal relationship between the auscultation features.
[0098] The migration unit converts the auscultatory feature sequence with the pathological acoustic causal intensity index into a causal state sequence for each disease mode, and performs disease diagnosis for the target individual through state migration between the causal state sequences.
[0099] If the above functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0100] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0101] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0102] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0103] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A big data diagnostic method based on auscultation data, characterized in that, include: The original auscultation sequence of the target individual is obtained using a stethoscope. The auscultation sequence is preprocessed, and pathologically sensitive segments within local time windows are screened based on multi-scale chaos measures to obtain segment sequences. The segment sequences are then feature-encoded using an encoder to obtain auscultation feature sequences. A temporal topology is constructed based on these auscultation feature sequences, and the pathological acoustic causal intensity index representing disease relevance in each auscultation feature is calculated by analyzing the causal relationships between the auscultation features. The auscultation feature sequences with the pathological acoustic causal intensity index are converted into causal state sequences for each disease modality, and disease diagnosis of the target individual is performed through state transitions between these causal state sequences.
2. The big data diagnostic method based on auscultation data as described in claim 1, characterized in that: The original auscultation sequence includes a sequence of target individual auscultation time-series signals continuously acquired by a stethoscope at a fixed sampling rate, reflecting the amplitude changes of heart sounds or lung sounds in the time dimension; the preprocessing includes amplitude normalization, pulse artifact removal, steady-state background noise suppression, and temporal boundary detection of the original auscultation sequence to construct a basic processing sequence for subsequent multi-scale chaotic measurement analysis.
3. The big data diagnostic method based on auscultation data as described in claim 2, characterized in that: The multi-scale chaos measure includes downsampling the original auscultation sequence based on n preset different downsampling ratios to generate n downsampling subsequences; the set of all sequences is denoted as: ; Using a sliding time window over time, the chaotic feature value of each sliding time window is calculated in both the original auscultation sequence and each of the downsampled subsequences; wherein, This represents a downsampled subsequence with a downsampling factor of n; on the time axis, according to arrive The process involves slicing time sequentially based on the chaotic feature values in each sequence. After the previous round of slicing is completed, the next round of slicing is performed on each slice. The slicing steps for each round are as follows: Step 1: Within any time period of the previous round's slice, extract the downsampled subsequence corresponding to the current round, and form a single-scale chaotic sequence representing the local complexity change of the sequence based on the chaotic feature values of each sliding time window. Step 2: In the single-scale chaotic sequence, determine the relative deviation of the chaotic feature values of adjacent windows according to the time sequence. If the relative deviation exceeds a preset threshold of the corresponding downsampling factor, then identify the end of the previous sliding window as a segmentation point. Step 3: Based on the segmentation point, perform segmentation in the previous round's slice. After segmentation, all slices of the current round are obtained. After completing the slicing of the current round, the next round of slicing is performed. After repeating n rounds, a fragment sequence of the original auscultation sequence is obtained.
4. The big data diagnostic method based on auscultation data as described in claim 3, characterized in that: The encoder includes inputting the time-domain signal from the slice into a neural network feature extraction structure, performing nonlinear mapping on the morphological features, frequency band energy distribution features, and short-time dynamic change features of each segment to obtain a multidimensional local feature vector of the slice; the neural network feature extraction structure includes a time-domain convolution branch that performs convolution operations on the original waveform, a time-frequency convolution branch that performs convolution operations on the time-frequency representation obtained by time-frequency transformation, and a time-series modeling branch for modeling the time dependencies within the segments; the time-domain convolution branch is used to extract the morphological features of the original waveform; the time-frequency convolution branch is used to extract the energy structure and frequency band distribution features; the time-series modeling branch is used to capture dynamic evolution features and extract short-time dynamic change features.
5. The big data diagnostic method based on auscultation data as described in claim 4, characterized in that: The temporal topology includes: treating slices as nodes, establishing edge relationships between any two nodes, and using a Bayesian network to generate probability distributions of the two nodes for each disease to represent the causal relationship between them; wherein, for any two nodes, the sum of probabilities of different diseases, health states, and unidentifiable conditions is in the range of 1; for any node in the temporal topology, after evaluating its ontology value, the corresponding slice is used as the starting point for evolution, and evolutionary analysis is performed in both positive and negative directions of the time axis, generating the overall value of the node based on the correlation strength of the nodes during the evolutionary analysis; if the ontology value or the overall value is lower than a corresponding threshold, the slice corresponding to the node is simulated by merging with the previous slice and the next slice, and the information gain of the ontology value and the overall value is calculated when merging forward and backward, selecting the merging direction with the goal of maximizing the weighted sum of the information gain, and merging slices according to the merging direction; the ontology value includes: for node i, generating a value coefficient based on the edge relationship between i and any node j. ;in, The standard deviation of different probabilities in the edge relationship between i and j is used to measure the unevenness of probabilities. Let represent the probability that the edge relationship between i and j is unidentifiable; then, by weighted summation of all edge relationships, we obtain the final ontology value of i. Where J represents the number of nodes; The weight between node i and node j is represented by an initial value A, which is proportionally decayed in both the forward and backward directions in the node sequence to generate a reference value for each node, which is then normalized. The evolutionary analysis includes inputting the node sequence and edge relationships into a pre-trained LSTM. The LSTM slides through the node sequence from front to back, and at each node, based on the edge relationships between the preceding nodes, generates probabilities of different diseases, health states, and unidentifiable conditions, resulting in a predicted probability distribution. The labels in the predicted probability distribution are used as dimensions, and the probability of each label is used as the value in the dimension to generate multi-dimensional features. The features of each dimension are fitted according to the order of the nodes. After calculating the relative deviation of node i in the fitting results of each dimension, the (1-relative deviation) of each dimension is summed to obtain the overall value of node i.
6. The big data diagnostic method based on auscultation data as described in claim 5, characterized in that: In the temporal topology structure after slicing and merging, the summation of the overall value of each node is divided by the number of nodes to obtain the pathological acoustic causality intensity index.
7. The big data diagnostic method based on auscultation data as described in claim 6, characterized in that: The causal state sequence includes the feature sequence of nodes in each dimension in the temporal topology structure after slicing and merging; the state transition includes learning the positive and negative correlations between different diseases based on prior probabilities; analyzing whether the changes in LSTM prediction probabilities between feature sequences of different dimensions conform to the positive and negative correlations: if a conforming correlation exists, then for diseases in the sequence whose LSTM prediction probabilities increase, the attention weights obtained after linear transformation according to the prior probabilities corresponding to the correlation are used for attention enhancement in LSTM to achieve the final output.
8. A big data diagnostic system based on auscultation data using the method described in any one of claims 1-7, characterized in that: The acquisition unit uses a stethoscope to acquire the original auscultation sequence of the target individual, preprocesses the auscultation sequence, and uses multi-scale chaos measure to screen pathologically sensitive segments in local time windows to obtain segment sequences; the encoding unit uses an encoder to encode the features of the segment sequences to obtain auscultation feature sequences. The calculation unit constructs a temporal topology based on the auscultation feature sequence, and calculates the pathoacoustic causal intensity index representing the disease correlation in each auscultation feature by analyzing the causal relationship between the auscultation features. The migration unit converts the auscultatory feature sequence with the pathological acoustic causal intensity index into a causal state sequence for each disease mode, and performs disease diagnosis for the target individual through state migration between the causal state sequences.
9. A computer device, comprising: A memory and a processor; the memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.