Respiration abnormity identification method, system and equipment and medium
By combining multi-channel convolutional networks and knowledge graphs, the problems of insufficient multimodal data fusion and lack of medical logical constraints are solved, high-accuracy recognition and adaptive detection of respiratory abnormalities are achieved, and the robustness and clinical interpretability of respiratory abnormality detection are improved.
Patent Information
- Application Number
- CN202510999331.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing respiratory abnormality detection technology suffers from insufficient multimodal data fusion, low accuracy in non-stationary signal feature extraction, and lack of medical logical constraints, resulting in high false detection rates and low clinical interpretability, making it difficult to achieve robustness and accuracy in complex environments.
Through multi-channel convolutional networks, multimodal data is cross-modally fused, combined with knowledge graphs and density clustering algorithms to extract features of audio, physiological movement and environmental monitoring data, and medical diagnostic rules are used to screen abnormal patterns to generate respiratory abnormality recognition results.
It improves the accuracy and clinical applicability of respiratory abnormality recognition, enhances the ability to process non-stationary signals, has strong adaptability, and can maintain stable recognition performance in complex environments.
Smart Images

Figure CN120753623A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical monitoring technology, and in particular to a method, system, device and medium for identifying abnormal breathing. Background Art
[0002] Existing respiratory anomaly detection technologies have significant limitations on multiple fronts. First, overreliance on a single signal source (e.g., using only breath sounds or motion signals) makes the system susceptible to environmental noise, resulting in a high false positive rate. Furthermore, it struggles to distinguish specific physiological activities (e.g., coughing) from pathological patterns, and its ability to capture events such as nocturnal apnea is insufficient. Second, multimodal data fusion often remains superficial, such as simply concatenating or weighted averaging independently processed signals. This fails to deeply explore complementary cross-modal correlations, leading to loss of signal correlations during sudden environmental changes and ineffectively capturing the complex nonlinear relationships between different physiological indicators (e.g., coughing and blood oxygen levels). Furthermore, traditional feature extraction methods (e.g., Fourier transforms) struggle to cope with the inherent nonstationary nature of respiratory signals and have low sensitivity to sudden events (e.g., sudden changes in the wheezing frequency band). Improved methods still have shortcomings in key areas such as component selection, leading to biased periodic feature extraction. Furthermore, most existing algorithms rely solely on data-driven approaches and lack the necessary clinical diagnostic rules. This can easily misclassify physiological changes (e.g., shortness of breath after exercise) as pathological abnormalities, reducing the clinical interpretability and reliability of the results. Finally, the application of medical knowledge bases such as knowledge graphs is currently limited to static disease classification. They fail to deeply integrate with feature clustering of real-time monitoring data, and are unable to dynamically identify and correct abnormal patterns that do not conform to known pathological mechanisms. These shortcomings collectively limit the robustness of the technology in complex environments, the depth of feature representation, the accuracy of clinical decision-making, and the system's adaptability. Summary of the Invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides a method for identifying abnormal breathing, which can solve the three core problems existing in existing abnormal breathing recognition technology: insufficient multimodal data fusion, low accuracy of non-stationary signal feature extraction, and lack of medical logic constraints.
[0005] To solve the above technical problems, the present application provides the following technical solutions, a respiratory abnormality recognition method, comprising: acquiring a multi-modal original data set containing audio signals, physiological motion signals and environmental monitoring data; extracting the spectral features of the audio signals, the periodic features of the physiological motion signals and the change features of the environmental monitoring data from the multi-modal original data set, and constructing a multi-dimensional feature matrix based on the extracted features; performing cross-modal fusion processing on the multi-dimensional feature matrix using a multi-channel convolutional network to generate a joint feature vector; combining a pre-constructed knowledge graph, identifying abnormal feature patterns from the joint feature vector through a density clustering algorithm, combining medical diagnosis rules to filter abnormal patterns, and outputting a respiratory abnormality recognition result.
[0006] As a preferred scheme of the respiratory abnormality recognition method described in the present application, wherein: the spectral features of the audio signals are decomposed into multiple frequency subbands by wavelet transform, and the energy distribution of each subband is calculated to obtain; The periodic features of the physiological motion signals are obtained by extracting the instantaneous frequency and instantaneous amplitude through Hilbert-Huang transform; The change features of the environmental monitoring data are obtained by extracting the time series statistical features and trend features through a sliding window.
[0007] As a preferred scheme of the respiratory abnormality recognition method described in the present application, wherein: the multi-channel convolutional network includes independent input channels corresponding to audio spectral features, physiological periodic features and environmental change features; Each channel uses a multi-scale convolution kernel group for local feature extraction, and generates a joint feature vector through feature splicing and dimension reduction operations.
[0008] As a preferred scheme of the respiratory abnormality recognition method described in the present application, wherein: the multi-scale convolution kernel includes a first scale convolution kernel for capturing local details, a second scale convolution kernel for extracting medium-range correlations, and a third scale convolution kernel for modeling global context relationships.
[0009] As a preferred scheme of the respiratory abnormality recognition method described in the present application, wherein: the density clustering algorithm dynamically assigns weights based on feature information entropy, generates a weighted feature space, calculates the information entropy of each dimension feature in the joint feature vector, and dynamically assigns weights according to the entropy value; Element-wise multiplication of the weights and the original feature matrix generates a weighted feature space; The medical diagnosis rule semantically associates the clustering results with the medical entities in the knowledge graph through the graph attention network, and uses the rule engine to exclude abnormal patterns that do not conform to clinical logic. The initial cluster center generated by density clustering is inserted into the knowledge graph as a virtual node, and the attention weight is calculated with the medical entity node; the rule engine matches the clinical diagnosis logic and screens out clusters that are inconsistent with the medical entity association.
[0010] As a preferred embodiment of the method for identifying abnormal breathing according to the present invention, the information entropy is calculated as follows: , in, represents the information entropy of the j-th dimension feature, n represents the number of analysis samples, Indicates the value of the i-th sample on the j-th dimension feature, and Respectively represent the maximum and minimum values of the j-th dimension feature in all samples, represents the smoothing constant.
[0011] As a preferred embodiment of the method for identifying abnormal breathing according to the present invention, the generating of the weighted feature space includes: , in, The weighted dimension is The feature space matrix of Represents the total number of dimensions of the feature, X represents the element in the i-th row and j-th column. , represents the element-wise multiplication operator, Indicates that the jth element is The feature weight vector of Represents the rank of the matrix.
[0012] The present invention provides a breathing abnormality recognition system.
[0013] To solve the above technical problems, the present invention provides the following technical solutions: a respiratory abnormality recognition system, which includes: a multimodal acquisition module, a cross-modal feature extraction module, a multi-channel convolution fusion module, and a knowledge graph clustering module; The multimodal acquisition module synchronously collects audio, physiological movement and environmental data through sensors and performs time alignment; The cross-modal feature extraction module extracts the spectral features of the audio signal; extracts the periodic features of the physiological signal; and extracts the temporal variation features of the environmental data. The multi-channel convolution fusion module integrates the features of different modalities into a multi-dimensional matrix and uses a multi-scale convolutional network to fuse cross-modal information to generate a joint feature vector; The knowledge graph clustering module combines the knowledge graph and density clustering algorithm to analyze abnormal patterns in feature vectors and screen credible results through medical rules.
[0014] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a breathing abnormality identification method when executing the computer program.
[0015] The present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of a method for identifying abnormal breathing are implemented.
[0016] The beneficial effects of the present invention are as follows: through the collaborative acquisition of multimodal data and deep feature fusion technology, combined with the dynamic constraint mechanism of the medical knowledge graph, the accuracy and clinical applicability of respiratory abnormality recognition are improved. The multi-scale feature extraction system constructed by wavelet transform and Hilbert-Huang transform effectively captures the transient frequency domain characteristics of respiratory sound signals and the time-varying periodic characteristics of physiological motion signals, overcoming the defect of insufficient analysis ability of non-stationary signals. Through the cross-modal fusion mechanism of the multi-channel convolutional network, deep correlation modeling of audio, physiological and environmental features in the time and space dimensions is achieved, and the characterization ability of abnormal patterns is enhanced. Combined with the density clustering algorithm driven by the knowledge graph, the entropy weight method is used to dynamically optimize the feature space weight distribution, and the clustering results are semantically corrected based on clinical diagnostic rules, so that the recognition results are consistent with the data distribution law and meet the medical logic requirements. The adaptive parameter adjustment mechanism effectively solves the problem of poor adaptability of traditional clustering algorithms to dynamic respiratory patterns. Through the synergistic effect of multi-scale convolution kernel groups and feature projection technology, the recognition of global abnormal patterns is enhanced while retaining local detail features. Through end-to-end feature optimization and knowledge fusion mechanism, it can maintain stable recognition performance under complex environmental noise interference, providing reliable technical support for early warning of respiratory diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A flowchart of a method for identifying abnormal breathing is provided in accordance with an embodiment of the present invention.
[0019] Figure 2 A schematic diagram of modules of a respiratory abnormality recognition system provided by one embodiment of the present invention.
[0020] In the figure: 11. Multimodal acquisition module; 12. Cross-modal feature extraction module; 13. Multi-channel convolution fusion module; 14. Knowledge graph clustering module. DETAILED DESCRIPTION
[0021] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0022] Example 1, with reference to Figure 1 , which is the first embodiment of the present invention, provides a method for identifying abnormal breathing, comprising: S1: Obtain a multimodal raw dataset containing audio signals, physiological motion signals, and environmental monitoring data.
[0023] Among them, the respiratory-related sound signals collected by acoustic sensors such as microphones include direct acoustic representations of respiratory abnormalities such as respiratory sounds (such as normal respiratory sounds, dry and wet rales, and wheezing sounds), coughs, and nasal sounds. Accelerometers, pressure sensors, or inertial measurement units are used to collect physiological motion data such as chest and abdominal movements, diaphragm fluctuations, and respiratory rate to reflect the amplitude, frequency, and coordination of respiratory movements. Temperature and humidity sensors, air pressure sensors, and gas concentration sensors (such as 、 Environmental parameters collected by devices such as sensors include temperature, humidity, air pressure, air quality (such as PM2.5), and oxygen concentration, which may induce or aggravate respiratory abnormalities. Portable devices (such as wearable bracelets and chest strap monitors) or fixed monitoring systems (such as hospital monitors) with integrated microphones, accelerometers, and temperature and humidity sensors can be used to achieve synchronous multi-signal acquisition. Through unified clock synchronization or timestamp marking, strict alignment of audio, physiological movement, and environmental data in the time dimension (such as a unified sampling frequency of 100Hz) is ensured, providing a temporal consistency foundation for subsequent cross-modal feature fusion.
[0024] S2: Extract the spectral features of audio signals, the periodic features of physiological motion signals, and the variation features of environmental monitoring data from the multimodal original data set, and construct a multidimensional feature matrix based on the extracted features.
[0025] For audio signals, discrete wavelet transforms (DWTs) or continuous wavelet transforms (CWTs) can be used to decompose the time-domain audio signal into different frequency subbands (e.g., a low-frequency respiratory baseband and a high-frequency abnormal respiratory sound band). The time-frequency localization of wavelet basis functions is exploited to capture the transient frequency components of respiratory sounds. Each audio signal sample generates a dimensional feature corresponding to the number of decomposition levels and subbands, reflecting the energy distribution differences across frequency bands. For physiological motion signals, empirical mode decomposition (EMD) can be used to adaptively decompose physiological motion signals (e.g., chest and abdominal acceleration signals) into several intrinsic mode functions (IMFs), each corresponding to periodic components at different time scales. For each IMF, the instantaneous frequency and instantaneous amplitude are extracted. The instantaneous frequency, defined as the derivative of the phase function, reflects the cyclical rate of change of respiratory motion, while the instantaneous amplitude reflects the time-varying characteristics of the motion amplitude. For environmental monitoring data, time series analysis can be performed. By segmenting environmental parameters (such as temperature, humidity, and air pressure) using a sliding window (window size w, step size Δt), statistical and trend features within each window are extracted to reflect the trend and fluctuation characteristics of the environmental parameters. Through modal alignment and dimensional unification, each row of the matrix corresponds to the feature space of a specific mode, and cross-modal dimension alignment is achieved through feature operations in the column direction to construct a multidimensional feature matrix including time-frequency domain, periodicity, and dynamic trends.
[0026] S3: Use a multi-channel convolutional network to perform cross-modal fusion processing on the multi-dimensional feature matrix to generate a joint feature vector.
[0027] The multi-channel convolutional network has multiple independent input channels, each corresponding to a modality in a multidimensional feature matrix—namely, audio spectrum features, physiological motion cycle features, and environmental monitoring change features. These features are input to the network through different channels. For each channel, a convolution kernel performs a convolution operation on the feature matrix to extract local features. For example, in the audio spectrum feature channel, the convolution kernel can capture local energy distribution patterns between different frequency subbands; in the physiological motion cycle feature channel, it can extract local changes in the respiratory cycle; and in the environmental monitoring change feature channel, it can detect local trends in environmental parameters. The convolution results from different channels interact through specific fusion mechanisms (such as feature concatenation and weighted summation) to achieve preliminary cross-modal feature fusion. Maximum pooling or average pooling is used. The features output by the pooling layer are integrated, and the high-dimensional features are mapped to a low-dimensional space through a fully connected layer, ultimately generating a joint feature vector.
[0028] S4: Combined with the pre-built knowledge graph, the density clustering algorithm is used to identify abnormal feature patterns from the joint feature vector and generate respiratory abnormality recognition results.
[0029] The pre-constructed knowledge graph is composed of medical professional knowledge, clinical data, and diagnosis experience of respiratory system diseases, and contains the association between various respiratory abnormal types (such as asthma, chronic obstructive pulmonary disease, pneumonia, etc.) and corresponding feature patterns, and the semantic association between different symptoms, signs, examination results and other medical entities. The density clustering algorithm is a clustering method based on the density distribution of data, which can automatically discover the clustering structure in the sample data, divide the sample points in the high-density area into the same clustering cluster, and divide the low-density area into the boundary between the clustering clusters. It has no strict restriction on the distribution form of data and is suitable for data with complex distribution patterns. The sample points in the joint feature vector are clustered by using the density clustering algorithm to obtain the initial clustering result, which may contain normal respiratory feature patterns and various abnormal respiratory feature patterns. The initial clustering result is input into the pre-constructed respiratory abnormal knowledge graph. The semantic association between the clustering cluster and the medical entity can be extracted by the graph attention network, and the clustering cluster can be screened according to the clinical diagnosis logic by using the rule engine to exclude abnormal patterns that do not conform to the diagnosis logic. A candidate abnormal feature set is obtained, and the matching degree of the candidate abnormal feature set with the pre-defined abnormal types in the knowledge graph is calculated by using a multilayer perceptron classifier to generate a respiratory abnormality recognition result, which provides a basis for the reference and decision of the result in actual application for doctors or users.
[0030] The above-mentioned respiratory abnormality recognition method provides support for respiratory abnormality recognition by obtaining a multi-modal original data set containing audio signals, physiological motion signals and environmental monitoring data from multi-dimensional data. The spectral features of the audio signals, the periodic features of the physiological motion signals and the change features of the environmental monitoring data are extracted from the data set to construct a multi-dimensional feature matrix, which can more fully excavate the deep association between different signals and effectively extract cross-modal complementary features. A multi-channel convolutional network is used to perform cross-modal fusion processing on the multi-dimensional feature matrix to generate a joint feature vector. In view of the dynamic changes of the respiratory signal, the internal relationship between the audio, physiological motion and environmental data is adaptively captured by the network model to solve the problem of weak processing ability for non-stationary characteristics of the signal and easy omission of key abnormal features. In combination with the pre-constructed knowledge graph, abnormal feature patterns are identified from the joint feature vector by using a density clustering algorithm to generate a respiratory abnormality recognition result. The recognition result is constrained and corrected by medical knowledge to exclude situations that do not conform to the clinical diagnosis logic, improve the accuracy and reliability of respiratory abnormality recognition, and meet the needs of actual medical scenarios.
[0031] In one of the embodiments, the spectral features of the audio signals, the periodic features of the physiological motion signals and the change features of the environmental monitoring data are extracted from the multi-modal original data set, and a multi-dimensional feature matrix is constructed based on the extracted features, including: S201, wavelet transform is performed on the audio signal to decompose it into sub-bands of different frequencies, and the energy distribution of each sub-band is extracted as the spectral feature of the audio signal; S202, processing the physiological motion signal using Hilbert-Huang transform, extracting the instantaneous frequency and instantaneous amplitude, and calculating the periodic characteristics of the physiological motion signal; S203, processing the environmental monitoring data using a time series analysis method, extracting parameter change trends and fluctuation characteristics of the data as change characteristics; S204 , arranging the extracted spectral features of the audio signal, the periodic features of the physiological motion signal, and the variation features of the environmental detection data in a preset order to obtain a multi-dimensional feature matrix.
[0032] Specifically, the low-frequency subband of an audio signal may reflect the fundamental frequency characteristics of breathing, while the high-frequency subband is more likely to capture specific signals such as abnormal breathing sounds. Using a wavelet transform to extract the energy distribution of each subband, the audio signal can be converted into a spectral signature that reflects the energy differences between different frequency components. The Hilbert-Huang transform is suitable for processing nonlinear and non-stationary signals, and physiological motion signals (such as those generated by chest and abdominal movement) possess such characteristics. This transform can be used to obtain the instantaneous frequency and amplitude of the signal. The instantaneous frequency reflects the speed of change in the respiratory movement cycle, such as an increase in instantaneous frequency during rapid breathing; the instantaneous amplitude reflects the magnitude of the movement. Environmental factors have a significant impact on respiratory health, and time series analysis methods are often used to process time-varying data. For environmental monitoring data (such as temperature, humidity, and air pressure), segmented processing using methods such as sliding windows can extract parameter change trends (such as a continuous rise in temperature) and fluctuation characteristics (such as the magnitude of the fluctuation) for each data segment. This reflects the dynamic changes in environmental factors and helps analyze potential links between environmental factors and respiratory abnormalities. The extracted audio signal spectral features, physiological motion signal periodic features, and environmental monitoring data change features are arranged in a preset order to construct a multidimensional feature matrix. The features of multimodal data are integrated into a data structure. Each row represents the features of a modality (audio, physiological motion, environmental monitoring), and the dimensions of different modal features are aligned in the column direction. This can retain the characteristic information of each modal data, facilitate subsequent cross-modal fusion processing, and provide comprehensive and rich data feature support for respiratory abnormality identification.
[0033] In one embodiment, the calculation process of extracting the spectral features of the audio signal, the periodic features of the physiological motion signal, and the variation features of the environmental monitoring data from the multimodal original data set and constructing a multidimensional feature matrix based on the extracted features is as follows: S301, use the following formula to calculate the audio spectrum characteristics : , in, represents the j-th layer wavelet basis function, represents the complex wavelet coefficient of the corresponding subband, represents the length of the k-th time window, represents the starting time point of the kth time window, represents the time resolution parameter, Indicates the number of decomposition levels; S302, using the following formula to calculate the physiological cycle characteristics : , in, represents the nth-order intrinsic mode function, represents the instantaneous amplitude, represents the instantaneous frequency, represents the number of effective IMF components; S303, calculate the environmental change characteristics using the following formula : , in, represents the environmental parameters within the sliding window w, 、 and represents the adaptive weighting coefficient optimized by KL divergence, w represents the total number of sliding windows, Indicates variance calculation, t is time, represents the second derivative, is the spatial gradient of environmental parameters; S304, use the following formula to construct a multidimensional feature matrix : , Among them, the number of rows of the multidimensional feature matrix is 3, corresponding to the three modalities of audio / physiological / environment, and the number of columns is , 、 and denote Hadamard product, tensor outer product and feature space projection respectively, represents the wavelet energy spectrum of the i-th time window, is the IMF order, is the audio feature dimension, i=1…m; represents the Gaussian time window standard operator, represents the center point of the time window, represents the adaptive window width, is the i-th sampling time point in the window; represents the K-th order modal differential characteristic, is the nth intrinsic mode function of the kth order decomposition, is the instantaneous amplitude of the corresponding IMF; represents a nonlinear mapping function, represents the trainable weight matrix, represents the time domain derivative characteristics, is the instantaneous frequency; represents the combined feature of the second-order derivative and variance, is the Laplace operator; represents the environmental parameter encoder, represents the temperature (T) / humidity (H) cross weight, Represents the frequency domain feature enhancement module.
[0034] For example, the audio spectrum feature calculation is to pass the audio signal through the j-th layer wavelet basis function Decompose into different sub-bands and obtain the complex wavelet coefficients of the corresponding sub-bands ; In the time window The coefficients of each sub-band are squared and integrated to obtain the normalized audio spectrum characteristics. ; Decompose the physiological signal into n-order intrinsic mode functions through empirical mode decomposition and perform Hilbert transform to calculate the instantaneous amplitude Instantaneous frequency , and the instantaneous frequency change rate Perform amplitude normalization and weighted averaging to highlight the dominant cycle components and obtain physiological cycle characteristics ; The environmental parameters in the sliding window w are calculated according to the total number of sliding windows Split, calculate variance Measuring the degree of data dispersion, index items Suppress excessive gradient fluctuations and second-order derivatives The linear change rate of the reaction parameter and the weighting coefficient 、 and Adaptive weighting to calculate environmental change characteristics ; The audio spectrum features , physiological cycle characteristics and environmental change characteristics Input into the multidimensional feature matrix to form 3 rows corresponding to the audio / physiological / environmental trimodalities and the number of columns is Dynamically align the time windows or multi-dimensional feature matrices of each modality , for audio lines: use Hadamard product to convert wavelet energy With Gaussian time window Weighted enhancement of temporal locality; for physiological behavior, tensor outer product fuses features of different order modal differentials and nonlinear mapping functions Capture multi-scale periodic interactions; for the environment, feature space projection converts the environment features Mapping to environment parameter encoder , enhancing frequency domain information. The aforementioned formula integrates audio, physiological, and environmental information, reducing interference from single-modal noise. Features such as instantaneous frequency and second-order derivatives accurately reflect the time-varying patterns of respiratory-related signals. The multidimensional matrix structure is compatible with multi-scale convolutional networks, providing high-information-density input for subsequent cross-modal fusion.
[0035] In one embodiment, a multi-channel convolutional network is used to perform cross-modal fusion processing on a multi-dimensional feature matrix to generate a joint feature vector, including: S401: Based on the multi-dimensional feature matrix, the parallel convolution channels of the multi-channel convolutional network are used to perform convolution kernel sliding processing on the spectral characteristics of the audio signal, the periodic characteristics of the physiological motion signal, and the change characteristics of the environmental monitoring data to obtain preliminary fusion features of each modality; S402, based on the requirements for feature extraction at different scales, a multi-scale convolution kernel group is used to perform a hierarchical convolution operation on the preliminary fusion features to obtain multi-scale feature information containing local details and global correlations and a convolution output feature map; S403, based on the convolution output feature map, using a batch normalization layer and a ReLU activation function layer to normalize and nonlinearly map the multi-scale feature information to obtain an optimized high-dimensional feature representation; S404, based on the principle of multi-channel feature splicing, uses tensor connection operations to fuse high-dimensional feature representations, and performs feature space projection and dimensionality reduction processing through a fully connected layer to obtain a joint feature vector.
[0036] Specifically, the multidimensional feature matrix is divided into three sub-matrices by modality (e.g., audio spectrum A, physiological cycle P, and environmental change E), each corresponding to an independent convolution channel. Convolution kernels are used in the audio channel to capture the time-frequency patterns of the spectrum (e.g., the harmonic structure of snoring); in the physiological channel, convolution kernels are used to extract temporal variations in the respiratory cycle (e.g., pause-resume patterns); and in the environmental channel, convolution kernels are used to integrate temporal correlations of parameters such as temperature and humidity. Local feature correlations within the modality are retained, and each channel outputs a preliminary fused feature. A hierarchical convolution operation is performed simultaneously on each preliminary fused feature using convolution kernels of three different scales. For example, a 3×3 kernel is used to capture local details of the feature (e.g., sudden spikes in respiratory sounds); a 5×5 kernel is used to extract mid-range correlations (e.g., the complete waveform of a respiratory cycle); and a 7×7 kernel is used to model global context (e.g., temperature trends throughout the sleep environment). Output feature maps at different scales are concatenated channel-by-channel to form a multi-scale feature tensor, which is then batch-normalized using a batch normalization layer. The ReLU activation function introduces nonlinearity and performs mapping, learning the complex nonlinear relationships required for breath recognition (such as the nonlinear coupling between ambient temperature and breathing depth) to obtain an optimized high-dimensional feature representation. The high-dimensional features of each modality are fused and concatenated channel-wise, and projected into the feature space and subjected to dimensionality reduction using a fully connected layer to produce a joint feature vector. Parallel channels and multi-scale convolutions are used to fully exploit the unique information of each modality. Tensor connections and an attention mechanism capture deep dependencies between modalities (such as how changes in ambient humidity affect breathing sound characteristics) and strengthen cross-modal correlations. Dilated convolutions, batch normalization, and dimensionality reduction ensure accuracy while reducing computational cost.
[0037] In one embodiment, S501, a multi-channel convolutional network is used to perform cross-modal fusion processing on a multi-dimensional feature matrix using the following formula: , in, is the fusion feature matrix, represents parallel convolution channel processing, is the loop index inside the convolution kernel, represents the m-th modal feature matrix of the input (m=1 audio, m=2 physiological, m=3 environmental), Indicates that the mth modal channel uses the sth convolution kernel of size P×Q, represents a two-dimensional convolution operation with zero padding, represents a multi-scale convolution kernel group, is the dilated convolution, represents batch normalization, represents the activation function, and Represent the mean and variance of the current batch, respectively. and represents the learnable scaling and translation parameters, represents a digital stability constant, represents feature concatenation and dimensionality reduction function, represents the transposed weight matrix of the fully connected layer; represents a nonlinear mapping function, Indicates negative saturation parameter.
[0038] Specifically, parallel convolution channel processing: Indicates that for each modality m, a multi-scale convolution kernel is used Perform parallel convolution operations, and each convolution result is batch normalized and relu activation , for the same modality m, the results of convolution kernels of different scales are spliced along the channel dimension: , the multi-scale features processed by the three modalities (audio, physiological and environmental) are spliced along the channel to obtain , through the transposed weight matrix of the fully connected layer Project high-dimensional features into low-dimensional space and use The ELU activation function enhances the model's robustness to input scale changes and adaptively adjusts feature distribution. Multi-scale convolution extracts intra-modal features and combines them with cross-modal splicing and fusion. Combined with adaptive normalization and nonlinear mapping, this generates a high-information-density joint feature vector, providing robust input for subsequent knowledge graph-based anomaly detection.
[0039] In a feasible embodiment, cross-modal feature fusion can be achieved through an attention mechanism, specifically, by mapping the audio spectrum, physiological cycle, and environmental change features to a high-dimensional embedding space. Attention weight calculation: The correlation weights between different modal features are calculated through the self-attention mechanism. For example: for the interaction between audio and environmental features, the weight of the influence of humidity changes on the high-frequency components of respiratory sounds is learned; for the interaction between physiological and environmental features, the correlation weight of temperature fluctuations on respiratory frequency is captured. The weighted sum of each modal feature is performed according to the attention weight to generate a joint feature vector. The fused features are reduced in dimension and normalized through a fully connected layer.
[0040] In another feasible embodiment, cross-modal feature fusion can also be achieved through a graph neural network. Specifically, multimodal features (audio, physiological, and environmental) are modeled as graph nodes, and inter-modal relationships are modeled as edges. A graph convolutional layer (GCN) aggregates information from adjacent nodes. For example, audio nodes and physiological nodes share an edge related to the respiratory cycle; environmental nodes and physiological nodes share an edge related to temperature and humidity. Node features are updated using a gating mechanism (such as a GRU) to capture long-term cross-modal dependencies. Global pooling is performed on the graph node features to generate a joint feature vector.
[0041] In one embodiment, identifying abnormal feature patterns from the joint feature vector using a density clustering algorithm includes: S601, based on the feature information entropy of each dimension in the joint feature vector, using the entropy weight method to assign a dynamic weight to each feature to generate a weighted feature space; S602, based on the sample distribution density in the weighted feature space, an adaptive parameter adjustment algorithm is used to dynamically optimize the neighborhood radius and the minimum number of points of the density clustering algorithm, and density clustering is performed to obtain an initial clustering result; S603: Input the initial clustering results into the pre-built respiratory abnormality knowledge graph, extract the semantic associations between clusters and medical entities through the graph attention network, and use the rule engine to eliminate abnormal patterns that do not conform to clinical diagnostic logic to generate a revised set of candidate abnormal features; S604: Based on the candidate abnormal feature set, the multi-layer perceptron classifier is used to calculate the matching degree between the candidate abnormal feature set and the predefined abnormality type in the knowledge graph, and the respiratory abnormality classification result and confidence score are output.
[0042] For example, the information entropy of each feature dimension j is calculated , reflecting the degree of disorder of the data in this dimension: low entropy ( →0) indicates that the data distribution is concentrated and the feature discrimination is high (important); high entropy ( →1) Description: The data is evenly distributed, and the feature discrimination is low (secondary). According to the information entropy calculation weight distribution dynamic weight, the weight of low entropy feature is higher, and the weighted feature space is generated, the key feature is enlarged, and the noise feature is suppressed. The adaptive parameter adjustment algorithm dynamically optimizes the neighborhood radius and minimum point number of the density clustering algorithm, ensuring that the number of samples in the core point neighborhood is ≥ minPts (neighborhood density threshold); the boundary point is in the core point neighborhood but does not meet minPts itself; noise point: neither core point nor boundary point; according to the clustering result, the core point and its reachable boundary point form a cluster, excluding noise points, to get the initial clustering result. The initial clustering result is input into the pre-constructed respiratory abnormality knowledge graph as node embedding, and the clustering cluster center feature vector is inserted into the knowledge graph as a virtual node; medical entities, such as asthma, are represented by pre-training embedding. Calculate the attention weight of the cluster node and the medical entity node, and associate the entity with high attention weight with the cluster. According to the rule engine correction, some clusters are associated with asthma, but PM2.5 in the environment data is normal, some clusters are associated with pneumonia but have no fever signs, etc. The clusters that do not meet the clinical diagnosis logic are excluded, and the corrected candidate abnormal feature set is generated. Use the multilayer perceptron classifier to match the pre-defined abnormal types (such as asthma, COPD) in the knowledge graph to calculate the matching degree and calculate the confidence degree, and output the respiratory abnormality classification result and confidence score. Combine dynamic feature weighting, enhance the role of key features through entropy weight method, and enhance the robustness of clustering; adaptive parameter adjustment: overcome the limitations of fixed parameters in traditional density clustering, adapt to complex data distribution; knowledge graph constraint: combine medical logic to exclude unreasonable abnormal patterns and improve the reliability of the result; end-to-end classification: from multi-modal data fusion to abnormal identification, forming a closed-loop system.
[0043] In a feasible embodiment, the abnormal pattern recognition can be realized by the isolation forest, specifically, by randomly selecting features and split values through the isolation forest algorithm, multiple "isolation trees" are constructed. According to the path length of the sample in the tree, the abnormal score is calculated (the shorter the path, the higher the abnormal probability). The abnormal score is associated with the disease threshold in the knowledge graph. For example: if the abnormal score of a sample is higher than the asthma threshold, but the environmental PM2.5 is normal, the rule engine correction is triggered; combined with clinical rules (such as apnea accompanied by blood oxygen drop), the abnormal result is filtered twice. Output the corrected abnormal type and confidence.
[0044] In another feasible embodiment, abnormal pattern recognition can also be achieved through reconstruction anomaly detection based on a deep autoencoder. Specifically, the deep autoencoder is trained using normal breathing data to learn the reconstruction capabilities of multimodal features. The reconstruction error (such as mean squared error) of the test sample is calculated. The larger the error, the higher the probability of an anomaly. The reconstruction error is mapped to the disease pattern in the knowledge graph. For example, a high reconstruction error that matches the characteristic pattern of asthma is considered an asthma anomaly. A rule engine is then used to eliminate noisy samples with high error but no pathological association. The anomaly type and a confidence score based on the error distribution are output.
[0045] In one embodiment, based on the feature information entropy of each dimension in the joint feature vector, an entropy weight method is used to assign a dynamic weight to each feature to generate a weighted feature space, which is achieved through the following calculation steps: S701, calculate the feature information entropy using the following formula : , in, represents the information entropy of the j-th dimension feature, n represents the number of analysis samples, Indicates the value of the i-th sample on the j-th dimension feature, and Respectively represent the maximum and minimum values of the j-th dimension feature in all samples, represents the smoothing constant; S702: Calculate the entropy weight based on the feature information entropy using the following formula: , in, and represents the weight of the j-th dimension feature, Indicates the total number of dimensions of the feature, and Represent the information entropy of the j-th and k-th dimension features respectively; S703, generate a weighted feature space using the following formula: , in, The weighted dimension is The feature space matrix of X represents the element in row i and column j. , represents the element-wise multiplication operator, Indicates that the jth element is The feature weight vector of Represents a time window.
[0046] Specifically, the characteristic information entropy is calculated to measure the data distribution uncertainty of the audio spectrum, physiological cycle, and environmental change feature dimensions, and to identify key features with high discrimination. Scaling to the interval [0, 1] , smoothing constant Possible values are , used to avoid the denominator being 0, that is, when a feature is all the same value, it is considered an invalid feature. Calculate the entropy of the normalized value Reflects the degree of disorder of the feature, low entropy ( →0) indicates that the data distribution is concentrated and the feature discrimination is high (important). For example, a breathing audio feature is significantly higher in abnormal samples than in normal samples; high entropy ( →1) Explanation: The data is evenly distributed, and the feature discrimination is low (minor). For example, in a certain recognition, the ambient temperature fluctuates little among all samples, and it is impossible to distinguish abnormalities. The weight of each feature The complement of its information entropy The proportion of the total complement is determined by and Satisfy weight normalization, if →0 → , the weight is significant at this time, if →0 →0, the weight can be ignored. Each column (j-th dimension feature) is multiplied by its weight ,Right now , which enhances key features and suppresses redundant ones. It automatically identifies important features based on data distribution without manual weighting, adapting to different scenarios (such as different environments or populations). The weighted feature space highlights key feature differences, making density clustering more likely to detect true abnormal patterns. The influence of low-weight features (such as environmental noise) is suppressed, reducing false positives. It quantifies feature discrimination based on information entropy and generates a weighted feature space, enhancing the model's sensitivity to key features. Without manual intervention, it is suitable for efficient analysis of multimodal respiratory data.
[0047] Example 2, reference Figure 2 , is an embodiment of the present invention, providing a respiratory abnormality recognition system, comprising a multimodal acquisition module 11, a cross-modal feature extraction module 12, a multi-channel convolution fusion module 13, and a knowledge graph clustering module 14; The multimodal acquisition module 11 synchronously collects audio, physiological motion and environmental data through sensors and performs time alignment; The cross-modal feature extraction module 12 extracts the spectral features of the audio signal; extracts the periodic features of the physiological signal; and extracts the temporal variation features of the environmental data. The multi-channel convolution fusion module 13 integrates the features of different modalities into a multi-dimensional matrix and uses a multi-scale convolutional network to fuse cross-modal information to generate a joint feature vector; The knowledge graph clustering module 14 combines the knowledge graph and the density clustering algorithm to analyze abnormal patterns in the feature vectors and screen credible results through medical rules.
[0048] This embodiment also provides an electronic device suitable for a method for identifying abnormal breathing, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement a method for identifying abnormal breathing as proposed in the above embodiment.
[0049] This embodiment further provides a storage medium storing a computer program, which, when executed by a processor, implements a breathing abnormality identification method as proposed in the above embodiment.
[0050] The storage medium proposed in this embodiment and the method for implementing abnormal breathing recognition proposed in the above embodiment belong to the same inventive concept. Technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0051] From the above description of the embodiments, those skilled in the art will clearly understand that the present invention can be implemented using software and necessary general-purpose hardware. Of course, it can also be implemented using hardware, but in many cases the former is the preferred embodiment. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This software product can be stored on a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disk, and includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0052] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for identifying abnormal breathing, characterized by: include, Obtain a multimodal raw data set containing audio signals, physiological motion signals, and environmental monitoring data; Extract the spectral features of audio signals, the periodic features of physiological motion signals, and the variation features of environmental monitoring data from the multimodal raw data set, and construct a multidimensional feature matrix based on the extracted features; A multi-channel convolutional network is used to perform cross-modal fusion processing on the multi-dimensional feature matrix to generate a joint feature vector; Combined with the pre-built knowledge graph, the density clustering algorithm is used to identify abnormal feature patterns from the joint feature vector, and the abnormal patterns are screened in combination with medical diagnostic rules to output the respiratory abnormality recognition results.
2. The method for identifying abnormal breathing according to claim 1, wherein: The spectral characteristics of the audio signal are decomposed into multiple frequency sub-bands by wavelet transform, and the energy distribution of each sub-band is calculated; The periodic characteristics of the physiological motion signal are obtained by extracting the instantaneous frequency and instantaneous amplitude through Hilbert-Huang transform; The change characteristics of the environmental monitoring data are obtained by extracting time series statistical characteristics and trend characteristics through a sliding window.
3. The method for identifying abnormal breathing according to claim 2, wherein: The multi-channel convolutional network includes independent input channels corresponding to audio spectrum characteristics, physiological cycle characteristics and environmental change characteristics; Each channel uses a multi-scale convolution kernel group to extract local features, and generates a joint feature vector through feature splicing and dimensionality reduction operations.
4. The method for identifying abnormal breathing according to claim 3, wherein: The multi-scale convolution kernel includes a first-scale convolution kernel for capturing local details, a second-scale convolution kernel for extracting medium-range associations, and a third-scale convolution kernel for modeling global contextual relationships.
5. The method for identifying abnormal breathing according to claim 4, wherein: The density clustering algorithm dynamically assigns weights based on feature information entropy, generates a weighted feature space, calculates the information entropy of each dimensional feature in the joint feature vector, and dynamically assigns weights based on the entropy value; Perform element-wise multiplication of the weights with the original feature matrix to generate a weighted feature space; The medical diagnosis rule semantically associates the clustering results with the medical entities in the knowledge graph through the graph attention network, and uses the rule engine to exclude abnormal patterns that do not conform to clinical logic. The initial cluster center generated by density clustering is inserted into the knowledge graph as a virtual node, and the attention weight is calculated with the medical entity node; the rule engine matches the clinical diagnosis logic and screens out clusters that are inconsistent with the medical entity association.
6. The method for identifying abnormal breathing according to claim 5, wherein: The calculation formula of the information entropy is: , in, represents the information entropy of the j-th dimension feature, n represents the number of analysis samples, Indicates the value of the i-th sample on the j-th dimension feature, and Respectively represent the maximum and minimum values of the j-th dimension feature in all samples, represents the smoothing constant.
7. The method for identifying abnormal breathing according to claim 6, wherein: The generating of the weighted feature space includes: , in, The weighted dimension is The feature space matrix of Represents the total number of dimensions of the feature, X represents the element in the i-th row and j-th column. , represents the element-wise multiplication operator, Indicates that the jth element is The feature weight vector of Represents the rank of the matrix.
8. A breathing abnormality recognition system, using the breathing abnormality recognition method according to any one of claims 1 to 7, characterized in that: Including multimodal acquisition module, cross-modal feature extraction module, multi-channel convolution fusion module, and knowledge graph clustering module; The multimodal acquisition module synchronously collects audio, physiological movement and environmental data through sensors and performs time alignment; The cross-modal feature extraction module extracts the spectral features of the audio signal; extracts the periodic features of the physiological signal; and extracts the temporal variation features of the environmental data. The multi-channel convolution fusion module integrates the features of different modalities into a multi-dimensional matrix and uses a multi-scale convolutional network to fuse cross-modal information to generate a joint feature vector; The knowledge graph clustering module combines the knowledge graph and density clustering algorithm to analyze abnormal patterns in feature vectors and screen credible results through medical rules.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the breathing abnormality identification method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a breathing abnormality identification method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Remote breathing health monitoring and management system and method
CN121075711A
Remote respiratory health monitoring and management system and method
CN121075711B
Breathing condition detection method and device
CN121242544A