Internal arteriovenous fistula anomaly detection method and system based on reversible audio-video transformation and diffusion model

The method uses reversible spectrogram transformation and diffusion models to enhance the accuracy and robustness of AVF stenosis detection by filtering noise and identifying high-frequency anomalies, addressing the limitations of existing machine learning models in AVF stenosis detection.

CN120304789AActive Publication Date: 2025-07-15ZHEJIANG NORMAL UNIV

Patent Information

Application Number
CN202510807919.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In the prior art, when using internal fistula tremor to detect internal fistula, the model has low accuracy and poor anti-interference ability, making it difficult to effectively identify abnormal states under complex distributions.

Method used

Using a method based on reversible audio-visual transformation and diffusion model, the internal fistula tremor samples were collected, pre-processed and converted into amplitude spectrum and phase spectrum, the diffusion model was used for denoising and reconstruction, combined with the multi-cluster Gaussian clustering model for abnormal classification, and the adaptive filter noise reduction and high-frequency attention mechanism were used to improve the anti-interference ability of the model.

Benefits of technology

It significantly improves the accuracy and anti-interference ability of fistula stenosis status detection, can stably extract key tremor characteristics in different environments, avoid noise interference, and improves the robustness and accuracy of abnormal detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120304789A_ABST
    Figure CN120304789A_ABST
Patent Text Reader

Abstract

The invention relates to the field of medical health monitoring, and discloses an internal arteriovenous fistula anomaly detection method and system based on a reversible audio-video transformation and diffusion model, and the method comprises the steps: collecting an internal fistula tremor sound sample; preprocessing the internal fistula tremor sound sample; converting the preprocessed internal fistula tremor sound sample into an amplitude spectrogram and a phase spectrogram by using reversible audio-video conversion; inputting the amplitude spectrogram into a diffusion model trained by a normal sample, and generating an amplitude spectrogram matched with the normal sample distribution through denoising reconstruction; the reconstructed amplitude spectrogram is combined with the original phase spectrogram to be restored into a reconstructed internal fistula tremor sound sample through reversible audio-video transformation; multi-band feature extraction is performed on the original and reconstructed internal fistula tremor sound samples, a multi-cluster Gaussian clustering model is trained, anomaly classification is realized based on a log-likelihood threshold, and internal fistula stenosis is judged. According to the technical scheme, the accuracy and robustness of internal arteriovenous fistula anomaly detection can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical health monitoring, and particularly to a method and system for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion models. Background Art

[0002] An arteriovenous fistula is the lifeline of hemodialysis patients. The thrill sound of the fistula has advantages such as simplicity and non-invasiveness in reflecting the stenosis of the fistula. Currently, the detection and evaluation of fistula stenosis mostly rely on manual work, and the accuracy is highly restricted by the experience and skills of the detection personnel. There are problems such as high dependence on professional skills and strong subjectivity of evaluation criteria. Using machines to objectively and autonomously discriminate the thrill sound of the fistula is of great significance for early stenosis warning of the fistula.

[0003] Currently, domestic and foreign teams have conducted relevant research on using the thrill sound of the fistula to discriminate the stenosis of the fistula. For example, Patent CN118749925 discloses an arteriovenous fistula monitoring bracelet that uses a thrill sound sensor and a blood flow sound sensor in combination with a deep learning network for abnormal warning; for example, the team of Cornell University School of Medicine (npj Digital Medicine, (2023) 6:163) proposed converting the thrill sound of the fistula into a Mel spectrogram and combining it with a classic deep learning network to judge the stenosis degree of the fistula. However, the data distribution of arteriovenous fistulas is often relatively complex, with characteristics such as multi-modal and non-Gaussian. When traditional deep learning models model such complex distributions, they are easily interfered by environmental noise and the discrimination effect is poor. Therefore, improving the accuracy and anti-interference ability of fistula stenosis state detection has become an urgent technical problem to be solved. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a method and system for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion models, which can effectively solve the problems of low model accuracy and poor anti-interference ability existing in traditional fistula stenosis state abnormal detection algorithms, so as to improve the accuracy of fistula stenosis state detection.

[0005] In a first aspect, the present application provides a method for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion models, characterized in that the method includes the following steps:

[0006] Step S10, collecting thrill sound samples of the fistula and establishing a sample database;

[0007] Step S20, preprocessing the thrill sound samples of the fistula;

[0008] Step S30, using reversible audio-visual transformation to convert the preprocessed thrill sound samples of the fistula into an amplitude spectrogram and a phase spectrogram;

[0009] Step S40: Input the amplitude spectrogram into the diffusion model trained with normal samples, and generate an amplitude spectrogram that conforms to the normal sample distribution through denoising reconstruction.

[0010] Step S50: Combine the reconstructed amplitude spectrogram with the original phase spectrogram and use the invertible audio-video transformation to restore it to the reconstructed arteriovenous fistula thrill sound sample.

[0011] Step S60: Divide the original and reconstructed arteriovenous fistula thrill sound samples into frequency bands, extract the reconstruction error and energy features of the samples and standardize them. After standardization, train a multi-cluster Gaussian clustering model, and set a threshold based on the log-likelihood value to achieve anomaly classification and judge arteriovenous fistula stenosis.

[0012] Furthermore, the specific method for establishing the sample database in Step S10 includes:

[0013] Step S11: Collect arteriovenous fistula thrill sound signals using a high-sensitivity capacitive sensor.

[0014] Step S12: Classify and label the arteriovenous fistula thrill sound samples.

[0015] Step S13: Convert the classified and labeled arteriovenous fistula thrill sound samples into a unified format and store them in the sample database, and establish a sample management system.

[0016] Furthermore, the specific method for preprocessing the arteriovenous fistula thrill sound samples in Step S20 includes:

[0017] Step S21: Denoise the arteriovenous fistula thrill sound samples through an adaptive filter.

[0018] Step S22: Adjust the frequency band gain of the arteriovenous fistula thrill sound samples through short-time Fourier transform to enhance the signal.

[0019] Step S23: Normalize the arteriovenous fistula thrill sound samples through the min-max normalization method.

[0020] Step S24: Standardize the normalized arteriovenous fistula thrill sound samples and perform outlier correction operations.

[0021] Furthermore, the specific method for converting the arteriovenous fistula thrill sound samples into an amplitude spectrum and a phase spectrum using short-time Fourier transform in Step S30 includes:

[0022] Step S31: Use short-time Fourier transform to convert the standardized and outlier-processed arteriovenous fistula thrill sound samples into a complex frequency spectrum ;

[0023] Among them, the complex spectrogram can be separated into an amplitude spectrum Sum phase spectrum ;

[0024] Step S32: Perform frequency-domain cropping on the amplitude spectrum and phase spectrum separated from the complex spectrum diagram to obtain the truncated spectrum diagram and ; ;

[0025] Step S33: Perform bilinear interpolation resampling on the truncated amplitude spectrum diagram and phase spectrum diagram respectively to resample them into an amplitude spectrum diagram and phase spectrum diagram with a fixed dimension;

[0026] Step S34: Perform normalization processing on the obtained amplitude spectrum diagram and phase spectrum diagram to compress their value ranges to .

[0027] Furthermore, save the key parameters in the short-time Fourier transform process, including the selected window function type, window length, frame shift step, number of Fourier transform points, etc. At the same time, save the minimum and maximum values of the amplitude spectrum and phase spectrum after being fixed in dimension (i.e., ) for restoring the original spectrum information and restoring the time-domain audio signal when performing the inverse short-time Fourier transform operation.

[0028] Furthermore, the specific method included in step S40 for denoising and reconstructing the original amplitude spectrum diagram to generate a reconstructed amplitude spectrum diagram is as follows:

[0029] Step S41: Represent the normalized two-dimensional amplitude spectrum diagram obtained in the above step S34 as , and perform a forward diffusion process on it, defining a series of gradually increasing noise scale parameters ( ). Construct the conditional probability distribution , and through the formula recurrence method, an image at any time step can be sampled from , and the calculation formula is:

[0030] ,

[0031] where is the cumulative attenuation factor, and is the noise sampled from the standard normal distribution.

[0032] Further, divide the spectrogram into sub-regions in the frequency dimension , and apply different noise perturbation parameters to each sub-region to generate a noisy image , where a larger noise scale is set for the high-frequency region.

[0033] Step S42, construct a noise prediction network based on U-Net, input the noisy amplitude spectrogram , and output the noise residual estimation value .

[0034] Further, introduce a high-frequency attention mechanism. Let the feature map be , calculate the frequency value through , define the frequency weighting function , construct the attention weight tensor , and obtain the high-frequency enhanced feature map through to guide the model to focus on the high-frequency region.

[0035] Step S43, train the noise prediction network, use the amplitude spectrogram samples of the arteriovenous fistula thrill sound labeled "normal" for learning, and sample the input spectrogram according to the forward diffusion sampling time step to obtain the noisy spectrogram , and send and into the network to obtain the noise prediction result .

[0036] Specifically, the network is trained using the logarithmic mean square error loss function (log-MSE), and the calculation formula of the loss function is:

[0037] ,

[0038] where is the original amplitude spectrogram, is the reconstructed amplitude spectrogram finally predicted by the model, is the smoothing term to prevent zero values in the logarithmic calculation.

[0039] Step S44, use the deterministic sampling strategy in Denoising Diffusion Implicit Models (DDIM) to start from the initial Gaussian noise image and perform the inverse diffusion process to gradually recover the clear amplitude spectrogram , where the deterministic sampling formula is:

[0040] ,

[0041] where is the noise attenuation coefficient for the current diffusion step, is the cumulative retention factor, is the random perturbation of the standard normal distribution.

[0042] Further, the step S50 restores and reconstructs the arteriovenous fistula thrill sound sample by using the inverse short-time Fourier transform, and the specific method includes:

[0043] Step S51, perform an anti-normalization operation on the reconstructed amplitude spectrogram to restore its amplitude size in the original physical quantity space.

[0044] Step S52, fuse the anti-normalized amplitude spectrogram obtained in step S51 with the phase spectrogram recorded in step S33 through to construct a complex frequency spectrogram , where is the time frame index, is the frequency index, is the imaginary unit.

[0045] Step S53, perform an inverse short-time Fourier transform on the complex frequency spectrogram to restore the time-domain thrill sound signal. Specifically, the discrete inverse Fourier transform can be used to obtain the time-domain signal segment corresponding to each frame ( , is the number of Fourier transform points).

[0046] Further, the frame shift method is used to overlap and add (Overlap-Add, OLA) the signal segments of each frame to form a complete time-domain signal , where is the frame shift step, is the total number of frames, is the window function.

[0047] Specifically, the inverse short-time Fourier transform uses the same parameter settings as the forward short-time Fourier transform in step S30, such as the window function type, window length, frame shift step, and number of Fourier transform points, to ensure the reversibility and physical consistency of the forward and inverse transforms.

[0048] Further, the step S60 classifies the anomalies of the sound sample, and the specific method includes:

[0049] Step S61, divide the frequency bands of the original and reconstructed arteriovenous fistula thrill sounds, calculate the reconstruction error and energy characteristics of each frequency band respectively, and establish a feature vector ;

[0050] Step S62: Standardize the features of the large-scale normal samples, train a multi-component Gaussian mixture model to fit the normal distribution features, and set the anomaly detection threshold based on the principle. ;

[0051] Step S63: Calculate the log-likelihood value corresponding to the frequency band features of the sample to be detected, compare it with the anomaly detection threshold to determine whether it is abnormal, and output the result.

[0052] In a second aspect, the present application also provides an arteriovenous fistula anomaly detection system based on reversible audio-visual transformation and diffusion model, which is characterized in that the system includes:

[0053] A collection module that uses a capacitive sensor to collect the fistula tremor sound samples generated by blood flowing through the fistula, discriminates the stenosis state of the fistula by professionals and labels relevant information, and establishes a sample database;

[0054] A preprocessing module that performs preprocessing such as noise reduction, signal enhancement, and normalization on the fistula tremor sound samples;

[0055] A reversible audio-visual transformation module that uses the short-time Fourier transform and its inverse transform to convert the fistula tremor sound samples into complex frequency spectra, separates the amplitude spectrum and phase spectrum, combines the reconstructed amplitude spectrum after passing through the image reconstruction module with the original phase spectrum, and restores it to the reconstructed fistula tremor sound samples through the inverse short-time Fourier transform;

[0056] An image reconstruction module that adds noise to the amplitude spectrum through the forward diffusion process, defines a series of gradually increasing noise scale parameters, divides the spectrogram sub-regions in the frequency dimension direction and applies different noise perturbation parameters, and sets a larger noise scale in the high-frequency region; then uses a denoising network based on U-Net and embedded with a high-frequency attention mechanism for reverse denoising to generate a reconstructed amplitude spectrum that fits the normal sample distribution;

[0057] A detection module that decomposes the frequency bands of the original and reconstructed fistula tremor sound samples, extracts the reconstructed error and energy ratio features of each frequency band; uses a multi-component Gaussian mixture model to model the normal sample feature distribution, and sets the log-likelihood threshold based on the principle; determines whether it is abnormal by comparing the log-likelihood value of the detection sample with the threshold, and outputs the result.

[0058] The present invention provides an arteriovenous fistula anomaly detection method and system based on reversible audio-visual transformation and diffusion model, which has the following beneficial effects:

[0059] 1. The adaptive filter is used to denoise the sound samples. By dynamically adjusting the filtering parameters, the environmental noise components are efficiently removed, and targeted enhancement is applied to the high-frequency abnormal regions, making the subtle turbulent signals significantly prominent in the spectrogram to highlight the subtle features in the arteriovenous fistula tremor sound, so that the model can stably extract the key tremor sound features in different clinical or life scenarios, effectively improving the system's resistance to interference and overall robustness.

[0060] 2. The one-dimensional tremor sound signal is converted into a two-dimensional spectrogram through the short-time Fourier transform, presenting the time-frequency local features and frequency components of the audio in the form of a spatialized "image". The diffusion model can directly learn the energy distribution law of the time-frequency two-dimensionality in the normal spectrogram, and is sensitive to the abnormal frequency components or energy distribution deviations that persist throughout the audio segment, and is sensitive to abnormal patterns, avoiding the feature blurring problem caused by long-sequence dependencies in time-domain modeling. Compared with the feature aliasing problem that may occur due to information stacking within a fixed duration when directly modeling time-domain signals, the two-dimensional frequency-domain modeling realizes the feature decoupling of "time position - frequency component" by transforming the time-domain dynamics into the frequency-domain spatial distribution, avoiding abnormal features being masked by normal signals in the same frequency band, and significantly improving the capture accuracy of the model for consistent abnormal patterns throughout the audio segment.

[0061] 3. The diffusion model is used to reconstruct the amplitude spectrogram of normal samples. By introducing a high-frequency attention mechanism, the model is guided to focus on the abnormal features in the high-frequency region, and the logarithmic loss function is used to amplify the error weight of the low-amplitude high-frequency components, making the model pay more attention to the rare high-frequency abnormal patterns in normal samples, and accurately identifying the spectral differences with weak amplitudes but significant pathological meanings even under noise interference, effectively avoiding the missed detection problem caused by low signal amplitudes or noise masking, and significantly improving the robustness and accuracy of abnormal detection. Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 It is a schematic diagram of the overall process of an arteriovenous fistula anomaly detection method based on reversible audio-visual transformation and diffusion model provided in an embodiment of the present invention.

[0064] Figure 2 It is a schematic diagram of the overall structure of an arteriovenous fistula anomaly detection system based on reversible audio-visual transformation and diffusion model provided in an embodiment of the present invention. Detailed Embodiments

[0065] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. The exemplary embodiments of the present invention shown in the drawings should be understood that these embodiments are provided to better convey the scope of the present invention to those skilled in the art. It should be understood that the flowcharts shown in the drawings are only illustrative examples, not necessarily including all contents and operations / steps, nor necessarily executed in the described order.

[0066] An embodiment of the present application provides a method for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion models. Among them, the method for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion models can be applied to servers and embedded devices. Among them, the server can be an independent server or a server cluster.

[0067] The following will describe some embodiments of the present application in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0068] Referring to Figure 1 , the present application provides a method for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion models. The method specifically includes steps S10 to S60:

[0069] Step S10, collect the fistula thrill sound samples and establish a sample database;

[0070] Step S20, preprocess the fistula thrill sound samples;

[0071] Step S30, use reversible audio-visual transformation to convert the preprocessed fistula thrill sound samples into amplitude spectrograms and phase spectrograms;

[0072] Step S40, input the amplitude spectrogram into the diffusion model trained with normal samples, and generate an amplitude spectrogram that conforms to the normal sample distribution through denoising reconstruction;

[0073] Step S50, combine the reconstructed amplitude spectrogram with the original phase spectrogram and use reversible audio-visual transformation to restore it to the reconstructed fistula thrill sound sample;

[0074] Step S60, perform frequency band division on the original and reconstructed fistula thrill sound samples, extract the reconstruction error and energy characteristics of the samples and standardize them. After standardization, train a multi-cluster Gaussian clustering model, and set a threshold based on the log-likelihood value to achieve abnormal classification and judge fistula stenosis.

[0075] Based on Figure 1 the embodiments shown, in this embodiment, step S10 includes:

[0076] Step S11: Use a high-sensitivity capacitive sensor to collect the tremor signals generated by the blood flow through the arteriovenous fistula. During the collection, ensure that the sensor is closely attached to the arteriovenous fistula site to minimize external environmental interference as much as possible. For each patient, multiple signal collections need to be carried out at different time points within the treatment cycle to ensure the diversity of sample data.

[0077] Step S12: Have physicians with rich clinical experience in vascular access jointly discriminate the collected samples. Classify the samples into abnormal samples (the arteriovenous fistula has stenosis) and normal samples (the arteriovenous fistula has no stenosis or the stenosis degree is within the normal range). Among them, each sample needs to be labeled with information such as patient ID, collection time, and diagnosis conclusion. After the labeled information is double-checked and verified, it is entered into the system.

[0078] Step S13: Uniformly convert the original tremor signals into WAV format files sampled at 16bit / 4kHz. Use a MySQL relational database to store the sample data, and establish a metadata table containing the following fields: signal file path, patient basic information (gender, age), collection parameters (sensor model, gain setting), diagnosis label (normal / abnormal), etc. At the same time, establish a perfect database management system to facilitate subsequent querying, updating, and maintenance of the sample data, and provide support for subsequent data analysis and model training.

[0079] Based on Figure 1 the embodiment shown, in this embodiment, Step S20 includes:

[0080] Step S21: Denoise the arteriovenous fistula tremor sound samples through an adaptive filter. Among them, the output of the filter of the adaptive filter and the error signal The calculation formula is:

[0081] ,

[0082] ,

[0083] Among them, is the transpose of the filter weight coefficient vector, is the reference noise signal.

[0084] The gain vector , weight vector and the inverse matrix of the input signal correlation matrix are updated according to the following formulas:

[0085] ,

[0086] ,

[0087] ,

[0088] Among them, is the step size factor, is the forgetting factor, is the input signal vector, is the filter weight coefficient vector, is the reference noise signal, is the filter output. is the inverse matrix of the input signal correlation matrix, is the input signal vector, is the error signal.

[0089] By continuously iteratively updating the weight vector and the inverse of the covariance matrix , the sum of the squares of the error signal is minimized, thereby realizing the adaptive filtering of the internal fistula tremor sound signal. This method is applied to the real-time monitoring and analysis of the internal fistula tremor sound, which can improve the signal-to-noise ratio of the signal and extract more accurate features.

[0090] Step S22: Adopt a signal enhancement method based on time-frequency analysis to enhance the signal of the sample to highlight the subtle features in the internal fistula tremor sound, facilitating subsequent image conversion and analysis.

[0091] Furthermore, by performing a short-time Fourier transform on the denoised internal fistula tremor sound sample, the energy distribution of the signal at different times and frequencies is obtained. According to the energy concentration characteristic of the internal fistula tremor sound in a specific frequency band, the signal in this frequency band is adjusted in gain to increase its energy amplitude and make it more prominent in the time-frequency spectrum.

[0092] Specifically, let the known internal fistula tremor characteristic frequency band be , and the corresponding frequency range is . According to the preset gain factor , the time-frequency spectrum values in this frequency band are adjusted as follows:

[0093] ,

[0094] Among them, is the gain factor. In this embodiment, is set to 2, is the signal after denoising the time-frequency spectrum after performing a short-time Fourier transform, where represents the time frame index, represents the frequency index.

[0095] Furthermore, for the adjusted time-frequency spectrum Perform the inverse short-time Fourier transform to obtain the time-domain enhanced signal . The specific calculation formula is as follows:

[0096] ,

[0097] where, is the reconstructed time-domain enhanced signal, is the number of signal sampling points, is the number of time frames, is the frequency quantity, is the adjusted time-frequency spectrum.

[0098] Step S23, perform a normalization operation on the samples by the min-max normalization method.

[0099] Specifically, after the previous noise reduction and enhancement processing, there may still be significant differences in the signal amplitude, energy and other characteristics of different samples. Normalize the signal to unify the dimension and data range to adapt to the subsequent image processing flow.

[0100] Specifically, for the arteriovenous fistula thrill sound sample signal , let its maximum value be , the minimum value be , and the normalized signal be , and its calculation formula is:[[]]

[0101] ,

[0102] Specifically, through the above steps, the signal amplitude is normalized to the interval to ensure that different samples are converted under the same numerical scale, avoiding deviations in the image processing results caused by excessive differences in the original data, and ensuring the stability and comparability of the image processing process.

[0103] Step S24, represent the arteriovenous fistula thrill sound sample after preliminary preprocessing as , where is the number of samples. First, calculate the mean and the standard deviation of this sequence, and their calculation formulas are respectively:[[]]

[0104] ,

[0105] ,

[0106] Furthermore, according to the above mean and the standard deviation , perform a standardization process on the original time series to obtain the standardized sample .

[0107] Further, to eliminate the influence of extreme outliers, for the sample points that satisfy , their values are corrected to . Among them, is the sign function, that is, when the input value is greater than 0, the output is 1; when it is equal to 0, the output is 0; when it is less than 0, the output is -1. Among them, is the threshold. Specifically, in this embodiment, the threshold is set to 3.

[0108] Based on Figure 1 the embodiment shown, in this embodiment, step S30 includes:

[0109] Step S31, convert the time series after standardization and outlier processing into a frequency-domain representation. Specifically, perform frequency-domain analysis on using the short-time Fourier transform to obtain the complex spectrogram , and its calculation formula is:

[0110] ,

[0111] Among them, represents the window function, represents the time point at the center of the current frame, represents the frequency index, is the number of points of the Fourier transform, is the length of the window function.

[0112] Specifically, in the embodiment of the present invention, the window function is set to:

[0113] ,

[0114] Among them, is the discrete time index, is the length of the window function.

[0115] Specifically, in the embodiment of the present invention, the length of the window function is set to 512, and the number of points of the Fourier transform is set to 1024.

[0116] Further, the amplitude spectrum and the phase spectrum can be separated from the complex spectrogram .

[0117] Step S32, for the amplitude spectrum and the phase spectrum separated from the complex spectrogram Perform frequency-domain cropping. Since only the frequency-domain information in the frequency range of 0–8000 Hz is concerned in this embodiment, the sampling rate and the number of Fourier transform points set by the short-time Fourier transform according to the above step S32 are used to calculate the actual frequency corresponding to each frequency point. The calculation formula is:

[0118] , ,

[0119] where is the frequency-point index, and is the sampling rate. Specifically, in the embodiment of the present invention, the sampling rate is set to 16000 Hz.

[0120] Further, retain all frequency indices that satisfy Hz, and crop the amplitude spectrum and the phase spectrum to obtain the truncated spectrogram and .

[0121] Step S33: Perform bilinear interpolation resampling on the truncated amplitude spectrogram and the phase spectrogram respectively, so that they are resampled into an amplitude spectrogram with a fixed dimension and a phase spectrogram . Specifically, in the embodiment of the present invention, the fixed dimension is set to .

[0122] Specifically, assume that the size of the original spectrogram is , and the size of the target spectrogram is . For any pixel position in the target spectrogram, its corresponding floating-point coordinates in the original spectrogram are:

[0123] ,

[0124] Find its four adjacent pixels in the original spectrogram, and calculate the interpolation value . The calculation formula for the interpolation value is:

[0125] ,

[0126] where , , represents the interpolation value of the target pixel position, and represents the values of the four adjacent points in the original spectrogram.

[0127] Perform the above operations on all pixels to complete the resampling of the amplitude spectrum and the phase spectrum, and finally obtain amplitude spectrum images with a fixed dimension and phase spectrum images .

[0128] Step S34: Normalize the and obtained above so that their value ranges are compressed to , and obtain and . Specifically, the normalization calculation formula is as follows:

[0129] ,

[0130] ,

[0131] where , are respectively the minimum and maximum values of the amplitude spectrum before normalization, and , are respectively the minimum and maximum values of the phase spectrum before normalization.

[0132] Furthermore, to ensure the consistency and traceability of spectrum restoration during audio reconstruction, this embodiment also records the key parameters in the short-time Fourier transform process, including the selected window function type, window length, frame shift step, number of Fourier transform points, etc. At the same time, save the minimum and maximum values of the amplitude spectrum and phase spectrum after being fixed in dimension (i.e., ) for restoring the original spectrum information and restoring the time-domain audio signal when performing the inverse short-time Fourier transform operation.

[0133] Based on the Figure 1 illustrated embodiment, in this embodiment, step S40 reconstructs a two-dimensional amplitude spectrum image, including:

[0134] Step S41: Forward diffusion process. Represent the normalized two-dimensional amplitude spectrum image obtained in step S30 as , where and respectively represent the horizontal and vertical coordinates of the image, and the image size is . Define a series of gradually increasing noise scale parameters , where , where is the total number of diffusion steps. In this embodiment, the number of diffusion steps is taken as 1000.

[0135] Specifically, define the noise scale coefficient at each time step , construct the following conditional probability distribution to implement the noise addition process of the two-dimensional amplitude spectrum diagram:

[0136] ,

[0137] wherein, , represents the unit covariance matrix.

[0138] Furthermore, through the formula recurrence method, the diffusion process can directly sample the image at any time step from : :

[0139] ,

[0140] wherein, is the cumulative attenuation factor.

[0141] To enhance the model's response ability to abnormal high-frequency structures, in this embodiment, the spectrogram is divided into multiple sub-regions in the frequency dimension , and different noise perturbation parameters are applied to each sub-region to generate the noisy image :

[0142] ,

[0143] wherein, is the noise sampled from the standard normal distribution .

[0144] Specifically, a larger noise scale is set for the high-frequency region to enhance its perturbation sensitivity.

[0145] Step S42, in the reverse denoising process, it is necessary to denoise and reconstruct the noisy amplitude spectrogram generated at each diffusion step. For this purpose, a noise prediction network is constructed to predict the random noise component contained in the noisy image at each step. Its structure is based on U-Net, composed of a symmetric encoder and decoder, and a high-frequency attention mechanism is embedded in all skip connection layers, thereby enhancing the model's ability to model high-frequency tremor structures.

[0146] Specifically, the input of the noise prediction network is the noisy amplitude spectrogram at the current diffusion time step , and the output is the corresponding noise residual estimation value , wherein are the network parameters.

[0147] Furthermore, to enhance the network's feature modeling ability in the high-frequency structure region, a high-frequency attention mechanism is introduced in the U-Net network structure in this embodiment. This mechanism constructs a weighting factor on the frequency axis to guide the network to enhance its perception ability of high-frequency information, thereby more effectively identifying typical high-frequency abnormal patterns in tremor sounds such as blood flow abnormalities and turbulence characteristics.

[0148] Specifically, let the feature map at a certain skip connection in the network be a tensor , where represents the number of channels, represents the frequency axis dimension, represents the time axis dimension. The sampling rate is set to , then the actual frequency value corresponding to the -th frequency index is:

[0149] ,

[0150] In this embodiment, a frequency weighting function is also set to mark the attention intensity at each frequency position. Its definition is as follows:

[0151] ,

[0152] where is the high-frequency enhancement factor. In this embodiment, the value range of is . The above frequency weights are extended in the channel dimension and broadcast along the time dimension, thereby constructing an attention weight tensor with a shape of . This weight tensor performs an element-wise multiplication operation with the original feature map to obtain a high-frequency enhanced feature map , and its calculation method is:

[0153] ,

[0154] The enhancement operation does not change the structure and parameters of the network, but is only embedded as a modulation module in the feature propagation path, effectively guiding the model to strengthen its attention to the high-frequency region, and improving the representation ability of high-frequency abnormalities in the tremor audio spectrum without introducing additional complexity.

[0155] Step S43, training the noise prediction network so that it can accurately estimate the Gaussian noise components added in each diffusion step, thereby supporting the spectrogram reconstruction in the inverse diffusion process. Only the amplitude spectrum samples of the arteriovenous fistula tremor sound labeled as "normal" are used for supervised learning in the training process to capture the normal distribution characteristics in the spectrogram.

[0156] During the training phase, the input spectrogram is randomly sampled for the diffusion time step according to the forward diffusion process , and the noisy spectrogram at the -th step is obtained . The diffusion process satisfies the following form:

[0157] ,

[0158] where represents the noise image sampled from the standard normal distribution, is the cumulative retention factor, is the noise ratio of the diffusion process at the -th step

[0159] Subsequently, and the diffusion time step are fed into the network together, and the noise prediction result is output, which is used to restore the frequency domain structure of the original spectrogram

[0160] By minimizing the above loss function, the model is guided to learn the mapping relationship for recovering the noise residual from the noisy image, enabling it to have denoising ability, thereby achieving accurate reconstruction of the spectrogram

[0161] In this embodiment, to improve the sensitivity of the model to the high-frequency and low-amplitude regions in the tremor audio spectrogram, the logarithmic mean square error loss function (log-MSE) is used as the training objective to replace the conventional MSE loss function, so as to enhance the fitting ability of the model in the low-amplitude region, especially for high-frequency components. The definition of the log-MSE loss function is as follows:

[0162] ,

[0163] where is the original spectrogram, is the reconstructed spectrogram finally predicted by the model, is a smoothing term to prevent zero values in the logarithmic calculation, which is used to avoid instability problems when the input value is zero or extremely small. By minimizing the above log-MSE loss function, the network is guided to focus on the modeling accuracy of the low-amplitude region in the spectrogram, thereby improving the sensitivity to abnormal tremor features and providing more discriminative reconstruction results for the subsequent anomaly detection module

[0164] Step S44, after the model training is completed, using the deterministic sampling strategy in Denoising Diffusion Implicit Models (DDIM), starting from the initial Gaussian noise image to perform the inverse diffusion process, and gradually restoring the clear amplitude spectrogram ​

[0165] Specifically, let the initial random noise spectrogram be , and the diffusion time step decreases from to . At each time step, use the trained noise prediction network to perform denoising estimation on the current noisy spectrogram , calculate the corresponding residual , and update the spectrogram accordingly. The iterative calculation method is as follows:

[0166] ,

[0167] where is the noise attenuation coefficient of the current diffusion step, is the cumulative retention factor, and represents the random perturbation from the standard normal distribution.

[0168] Furthermore, in this embodiment, to accelerate the inference process and improve the reconstruction stability, an interval sampling step is set under the DDIM framework, that is, use less than steps for iteration to make the generation process converge efficiently to a stable spectrogram.

[0169] Specifically, the iterative process is executed sequentially from to , and finally a reconstructed spectrogram with the same size as the original amplitude spectrogram and the same frequency range is output.

[0170] Based on the embodiment shown in Figure 1 , in this embodiment, step S50 restores the tremolo sample based on the inverse short-time Fourier transform, including:

[0171] Step S51, perform denormalization processing on the reconstructed amplitude spectrogram described in step S44 to restore its amplitude size in the original physical quantity space. The calculation formula of the denormalization operation is as follows:

[0172] ,

[0173] where is the reconstructed amplitude spectrogram after denormalization, is the reconstructed amplitude spectrogram after normalization, and are respectively the minimum value and the maximum value of the amplitude spectrogram recorded in step S30 after being fixed in dimension.

[0174] Step S52, the amplitude spectrogram Fuse with the phase spectrogram after fixing the dimension described in step S33 to construct a reconstructed complex frequency spectrogram . Among them, represents the time frame index, represents the frequency index. The calculation method of the reconstructed complex frequency spectrogram is as follows:

[0175] ,

[0176] wherein, represents the imaginary unit, is the reconstructed complex frequency spectrogram, which contains the amplitude and phase information in the frequency domain.

[0177] Step S53: Perform an inverse short-time Fourier transform on the reconstructed complex frequency spectrogram to obtain the finally restored tremolo signal in the time domain , where n is the sampling point index.

[0178] Specifically, for the reconstructed complex frequency spectrogram described in step S52, through the discrete inverse Fourier transform, obtain the time domain signal segment corresponding to each frame , and its calculation formula is:

[0179] ,

[0180] wherein, is the number of Fourier transform points, represents the time frame corresponding signal segment, and the length is the real part sequence.

[0181] To ensure the continuity between frames and eliminate edge effects, the embodiments of the present invention adopt a frame shift method to superimpose and sum the signal segments of each frame to form a complete time domain signal , and its reconstruction calculation formula is:

[0182] ,

[0183] wherein, is the frame shift step size, is the total number of frames, is the window function, is the signal segment obtained by the inverse transform of the th frame. The sum results at all overlapping positions constitute the finally restored audio signal .

[0184] Further, to be consistent with the forward short-time Fourier transform operation in step S30, parameters such as the window function type, window function length, frame shift step, and number of Fourier transform points used in the inverse short-time Fourier transform are exactly the same as those in the short-time Fourier transform to ensure the reversibility and physical consistency of the forward and inverse transforms.

[0185] Based on Figure 1 the embodiments shown, in this embodiment, step S60 includes:

[0186] Step S61, perform multi-dimensional feature extraction and fusion on the original arteriovenous fistula thrill sound sample and its reconstructed arteriovenous fistula thrill sound sample To effectively characterize the reconstruction characteristics of the signal in different frequency bands, a finite impulse response (FIR) filter bank that satisfies the Nyquist sampling theorem is used to decompose the signal into frequency bands. The designed filter bank divides the signal into a low-frequency band (0–200 Hz), a middle-frequency band (200–800 Hz), and a high-frequency band (800–8000 Hz) through different cut-off frequencies. The transfer function of the th band filter is defined as:

[0187] ,

[0188] where are the filter coefficients corresponding to the respective frequency bands, is the filter order.

[0189] Further, calculate the root mean square error (RMSE) between the original signal and the reconstructed signal for each of the above three frequency bands to measure the fidelity of the reconstruction quality within the frequency band. The formula is as follows:

[0190] ,

[0191] where represents the set of sampling points corresponding to the th frequency band, is the number of points in this set.

[0192] Meanwhile, to reflect the energy distribution characteristics of the reconstructed signal in each frequency band, calculate its energy ratio:

[0193] ,

[0194] where is the total number of signal sampling points.

[0195] Further, combine the above 6 feature quantities into an original feature vector , and the Z-Score method is used to perform standardization processing based on the statistics of the normal samples in the training set to obtain the normalized feature vectors . Among them, the mean vector and the standard deviation vector used in the standardization process are respectively from the statistical results of the normal samples in the training set, ensuring that the features of each dimension have balanced weights in model training and avoiding training biases caused by dimensional differences.

[0196] In step S62, further based on the set of normalized normal sample feature vectors , the Expectation-Maximization (EM) algorithm is used to train a Gaussian mixture model (GMM) with 5 components to model the distribution of normal samples in the feature space. The probability density function of this model is:

[0197] ,

[0198] where is the weight of the th component (satisfying ), and are respectively the mean and covariance matrix of this component.

[0199] By iteratively maximizing the log-likelihood function of the training samples:

[0200] ,

[0201] the set of model parameters is optimized. Among them, in the E step of the EM algorithm, according to the current parameter estimation values, the posterior probability that each sample belongs to the th Gaussian component is calculated; in the M step, the samples are weighted according to the above responsibilities, and the parameter estimation values of each component are updated, including the mean, covariance, and weight coefficient.

[0202] Specifically, in the embodiment of the present invention, the number of components is set to 5, that is, 5 Gaussian components are used to model the feature distribution of normal samples.

[0203] After the model training is completed, the log-likelihood values of all normal samples are calculated, and then their mean and standard deviation

[0204] Furthermore, according to the 3σ principle, the anomaly detection threshold is set.

[0205] In step S63, for any sample to be measured, first extract its feature vector and normalize it to , and then calculate its log-likelihood value under the GMM model . If the log-likelihood value is lower than the threshold τ, then determine that the sample is abnormal; otherwise, it is normal.

[0206] Furthermore, to quantify the degree of abnormality, calculate its standardized log-likelihood ratio: ,

[0207] When , it indicates that the sample significantly deviates from the normal distribution, and the greater the degree of abnormality.

[0208] Furthermore, the final output result includes three items: classification label (normal / abnormal), abnormality index (Z value), and feature contribution degree. Among them, the feature contribution degree is obtained by calculating the absolute value of the gradient of the log-likelihood function with respect to each feature component , which is used to quantitatively measure the influence intensity of each frequency band feature on the classification result, so as to provide auxiliary information with interpretability and reference value for clinical practice.

[0209] Referring to Figure 2 , the present application also provides an arteriovenous fistula abnormality detection system based on reversible audio-visual transformation and diffusion model, which is characterized in that the system includes:

[0210] A collection module, which uses a capacitive sensor to collect the arteriovenous fistula tremor sound samples generated by blood flowing through the arteriovenous fistula, discriminates the stenosis state of the arteriovenous fistula by a professional and labels relevant information, and establishes a sample database;

[0211] A preprocessing module, which performs preprocessing such as noise reduction, signal enhancement, and normalization on the arteriovenous fistula tremor sound samples;

[0212] A reversible audio-visual transformation module, which uses the short-time Fourier transform and its inverse transform to convert the arteriovenous fistula tremor sound samples into complex frequency spectra, separates the amplitude spectrum and phase spectrum, and combines the reconstructed amplitude spectrum after passing through the image reconstruction module with the original phase spectrum, and restores it to the reconstructed arteriovenous fistula tremor sound samples through the inverse short-time Fourier transform;

[0213] An image reconstruction module, which adds noise to the amplitude spectrum through a forward diffusion process, defines a series of gradually increasing noise scale parameters, divides the spectrogram sub-regions in the frequency dimension direction and applies different noise perturbation parameters, and sets a larger noise scale in the high-frequency region; then uses a denoising network based on U-Net and embedded with a high-frequency attention mechanism to perform reverse denoising to generate a reconstructed amplitude spectrum that conforms to the normal sample distribution;

[0214] The detection module decomposes the original and reconstructed arteriovenous fistula thrill sound samples into frequency bands, extracts the reconstruction error and energy ratio characteristics of each frequency band; models the normal sample feature distribution using a multi-component Gaussian mixture model, and sets a logarithmic likelihood threshold based on the principle; determines whether it is abnormal by comparing the logarithmic likelihood value of the detection sample with the threshold, and outputs the result.

[0215] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A method for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion models, characterized in that, It includes the following steps: Step S10, collect the fistula thrill sound samples and establish a sample database; Step S20, preprocess the fistula thrill sound samples; Step S30, use the invertible audio - visual transformation to convert the preprocessed fistula thrill sound samples into an amplitude spectrogram and a phase spectrogram; Step S40, input the amplitude spectrogram into a diffusion model trained with normal samples, and generate an amplitude spectrogram that conforms to the normal sample distribution through denoising reconstruction; Step S50, combine the reconstructed amplitude spectrogram with the original phase spectrogram and use the invertible audio - visual transformation to restore it to a reconstructed fistula thrill sound sample; Step S60, calculate the reconstruction error and energy characteristics of the original and reconstructed fistula thrill sound samples in different frequency bands, form a feature vector, then train a multi - cluster Gaussian clustering model, and set a threshold based on the log - likelihood value to achieve abnormal classification and judge fistula stenosis.

2. The arteriovenous fistula abnormality detection method based on reversible audio-visual transformation and diffusion model according to claim 1, wherein The said step S20 includes the following steps: Step S21, perform noise reduction processing on the fistula thrill sound samples through an adaptive filter; Step S22, perform frequency - band gain on the fistula thrill sound samples through short - time Fourier transform to enhance the signal; Step S23, perform normalization processing on the fistula thrill sound samples through the min - max normalization method; Step S24, perform standardization processing on the normalized fistula thrill sound samples and perform outlier correction operations.

3. The arteriovenous fistula abnormality detection method based on reversible audio-visual transformation and diffusion model according to claim 2, wherein, The said step S22 performs frequency - band gain adjustment on the fistula thrill sound samples through short - time Fourier transform, and its calculation formula is: , Among them, is the spectrum of the adjusted arteriovenous fistula thrill sound, is the frequency corresponding to the characteristic frequency band of the arteriovenous fistula thrill, range, is the gain factor, is the arteriovenous fistula thrill sound signal after denoising, is the time-frequency spectrum after short-time Fourier transform, where represents the time frame index, represents the frequency, index; The calculation formula for obtaining the time - domain enhanced signal by performing inverse short - time Fourier transform on the adjusted time - frequency spectrum is: , Among them, is the enhanced signal of the reconstructed in-time-domain arteriovenous fistula thrill sound, is the number of sampling points of the arteriovenous fistula thrill sound signal, is the number of time frames, is the frequency quantity.

4. A method for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion model according to claim 1, characterized in that, The said step S30 includes the following steps: Step S31, use short - time Fourier transform to convert the standardized and outlier - processed fistula thrill sound samples into an amplitude spectrogram and a phase spectrogram; Step S32, perform frequency - domain cropping on the amplitude spectrogram and the phase spectrogram; Step S33, perform bilinear interpolation resampling on the truncated amplitude spectrogram and phase spectrogram respectively; Step S34, perform normalization processing on the resampled amplitude spectrogram and phase spectrogram.

5. The arteriovenous fistula abnormality detection method based on reversible audio-visual transformation and diffusion model according to claim 1, characterized in that, The said step S40 includes the following steps: Step S41: Define a gradually increasing noise scale parameter for the original amplitude spectrum diagram , construct the conditional probability distribution and recursively sample the noisy image through the formula, apply different noise perturbations in the frequency dimension by dividing regions, and set a larger noise scale in the high-frequency region; Step S42, construct a noise prediction network based on U - Net, introduce a high - frequency attention mechanism, construct an attention weight tensor through frequency value calculation and frequency - weighted function, and enhance the model's attention to the high - frequency region; Step S43, use the amplitude spectrogram samples labeled as "normal" for training, generate a noisy spectrogram through forward diffusion and input it into the network, and train the noise prediction network using the logarithmic mean square error loss function; Step S44, use the deterministic sampling strategy, start from the initial Gaussian noise image and perform the inverse diffusion process, and gradually restore the clear reconstructed amplitude spectrogram through the deterministic sampling formula.

6. The arteriovenous fistula abnormality detection method based on the reversible audio-visual transformation and diffusion model according to claim 5, wherein, The said logarithmic mean square error loss function is: , Among them, is the original amplitude spectrum diagram, is the reconstructed amplitude spectrum diagram finally predicted by the model, is a smoothing term to prevent zero values in logarithmic calculations.

7. A method for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion model according to claim 5, characterized in that, The said deterministic sampling formula is: , Among them, is the noise attenuation coefficient of the current diffusion step, is the cumulative retention factor, represents the random perturbation from the standard normal distribution.

8. The arteriovenous fistula abnormality detection method based on the reversible audio-visual transformation and diffusion model according to claim 1, characterized in that, The said step S50 includes the following steps: Step S51, perform anti - normalization processing on the reconstructed amplitude spectrogram to restore the original amplitude size; Step S52, fuse the anti - normalized reconstructed amplitude spectrogram with the resampled phase spectrogram to construct a reconstructed complex frequency spectrum diagram containing amplitude and phase information; Step S53: Perform an inverse short-time Fourier transform on the complex spectrogram, generate a complete time-domain reconstructed audio signal through discrete inverse Fourier transform and frame-shifted overlapping summation, and use the same inverse short-time Fourier transform parameters as the forward short-time Fourier transform to ensure reversibility.

9. A method for detecting arteriovenous fistula abnormalities based on reversible audio-visual transformation and diffusion model according to claim 1, characterized in that, The step S60 includes the following steps: Step S61: Divide the frequency bands of the original and reconstructed arteriovenous fistula thrill sounds, calculate the reconstruction error and energy characteristics of each frequency band respectively, and establish a feature vector ; Step S62: Standardize the features of the large-scale normal samples, train a multi-component Gaussian mixture model to fit the normal distribution features, and set the anomaly detection threshold based on the 3σ principle ; Step S63, calculate the log-likelihood value corresponding to the frequency band feature of the sample to be detected, and compare it with the anomaly detection threshold to determine whether it is abnormal and output the result.

10. A arteriovenous fistula abnormality detection system based on reversible audio-visual transformation and diffusion model, which is used to execute an arteriovenous fistula abnormality detection method based on reversible audio-visual transformation and diffusion model as described in any one of claims 1 to 9, and is characterized in that, It includes: An acquisition module, which is used to acquire samples of arteriovenous fistula thrill sounds and establish a sample database containing annotation information; A preprocessing module, which is used to perform preprocessing operations such as noise reduction, signal enhancement, and normalization on the samples of arteriovenous fistula thrill sounds; A reversible audio-visual transformation module, which is used to convert the samples of arteriovenous fistula thrill sounds into complex spectrograms, separate the amplitude spectrogram and the phase spectrogram, combine the reconstructed amplitude spectrogram generated after passing through the image reconstruction module with the original phase spectrogram, and restore it to the reconstructed samples of arteriovenous fistula thrill sounds through inverse transformation; An image reconstruction module, which is used to apply multi-scale noise perturbations to the amplitude spectrogram, perform step-by-step image reconstruction based on the diffusion model, introduce an attention mechanism to enhance the high-frequency region modeling ability, and generate a reconstructed spectrogram that fits the normal sample distribution; A detection module, which is used to perform frequency division analysis and feature extraction on the original and reconstructed thrill sound samples, construct a feature vector of frequency band error and energy ratio, train an abnormal discrimination model and set a discrimination threshold to realize the abnormal recognition and result output of the arteriovenous fistula stenosis state.

Citation Information

Patent Citations

  • Arterial aneurysm classification and recognition device, method, equipment and storage medium

    CN116523914A

  • Urine sample hyperspectral image-based glomerular disease classification diagnosis model construction method

    CN117036269A

  • Arteriovenous fistula three-dimensional model construction device and method

    CN120000250A

  • Internal fistula monitoring health management bracelet and system

    CN120048559A

  • Internal vascular fistula cuff with blood flow monitoring function and alarm function

    CN211355439U

Cited By

  • Method for extracting internal arteriovenous fistula blood flow sound signal features and electronic equipment

    CN120748456A

  • Abnormal radio signal monitoring method and system based on artificial intelligence

    CN121117904A