Method and device for predicting dialysis access narrowing using convolutional neural networks

A CNN-based method accurately predicts dialysis access narrowing using audio data, addressing the subjectivity of traditional auscultation and palpation methods by providing objective narrowing degree predictions.

JP7745652B2Active Publication Date: 2025-09-29IND ACADEMIC COOP FOUND YONSEI UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023569940
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-13
Filing Date
2022-05-13
Publication Date
2025-09-29
Estimated Expiration
2042-05-13

Smart Images

  • Figure 0007745652000022
    Figure 0007745652000022
  • Figure 0007745652000023
    Figure 0007745652000023
  • Figure 0007745652000024
    Figure 0007745652000024
Patent Text Reader

Abstract

A method and apparatus for predicting narrowing of a dialysis access route using a convolutional neural network according to a preferred embodiment of the present invention is capable of predicting the degree of narrowing of a dialysis access route of a subject from audio data of the dialysis access route of the subject based on a narrowing prediction model including a convolutional neural network (CNN), thereby more accurately predicting the degree of narrowing of the dialysis access route and thereby guiding further examinations and treatments.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for predicting dialysis access narrowing using a convolutional neural network, and more particularly to a method and apparatus for diagnosing dialysis access narrowings that require treatment. [Background technology]

[0002] Confirming the presence or absence of abnormalities in dialysis access routes, such as arteriovenous fistulas, relies heavily on palpation and auscultation. When palpating along the anterior and posterior sides of the narrowing, the thrill and pulsation felt vary greatly depending on the location. Vibrations felt during palpation can sometimes sound similar to high-pitched bruits in the audible frequency range when using a stethoscope. Similarly, the presence and intensity of bruits can be used to indirectly diagnose narrowing or occlusion of an arteriovenous fistula. However, there are few physicians skilled in auscultation, and many subjective factors influence judgments based on auscultation. Therefore, it is difficult to objectively distinguish between significant narrowings requiring treatment, such as angioplasty. Summary of the Invention [Problem to be solved by the invention]

[0003] An object of the present invention is to provide a method and apparatus for predicting the degree of narrowing of a dialysis access route using a convolutional neural network (CNN), which predicts the degree of narrowing of a dialysis access route of a subject from audio data of the dialysis access route based on a narrowing prediction model including a convolutional neural network (CNN).

[0004] Other objects of the present invention not explicitly stated can be further considered within the scope that can be easily inferred from the following detailed description and its effects. [Means for solving the problem]

[0005] To achieve the above object, a method for predicting narrowing of a dialysis access route using a convolutional neural network according to a preferred embodiment of the present invention includes the steps of acquiring audio data related to a dialysis access route of a subject, and predicting the degree of narrowing corresponding to the audio data based on a narrowing prediction model including a pre-trained convolutional neural network (CNN).

[0006] Here, the audio data obtaining step may include preprocessing the audio data, and the narrowing degree predicting step may include inputting the preprocessed audio data into the narrowing prediction model and predicting a narrowing degree corresponding to the audio data based on an output value of the narrowing prediction model.

[0007] Here, the acquiring of the audio data may include acquiring the audio data of a predetermined section from the audio data, acquiring a spectrogram based on the audio data of the predetermined section, normalizing the acquired spectrogram, and adjusting a size of the normalized spectrogram.

[0008] Here, the method may further include a step of training the stenosis prediction model based on a training dataset including first audio data for the dialysis access route acquired before the procedure and second audio data for the dialysis access route acquired after the procedure.

[0009] Here, the narrowing prediction model may take a spectrogram as an input and output a narrowing degree value.

[0010] Here, the narrowed prediction model training step may include preprocessing the training dataset, setting the first audio data as a first ground truth label and the second audio data as a second ground truth label, and training the narrowed prediction model based on the preprocessed training dataset.

[0011] Here, the narrowed prediction model training step may include, for each of the audio data included in the training dataset, acquiring audio data of a predetermined section from the audio data, acquiring spectrograms based on the audio data of the predetermined section, normalizing the acquired spectrograms, horizontally shifting the normalized spectrograms to increase the number of the spectrograms, and adjusting the size of the increased spectrograms, thereby preprocessing the training dataset.

[0012] Here, the narrowing prediction model learning step may include dividing the preprocessed learning dataset into a training dataset, a tuning dataset, and a validation dataset according to a predetermined criterion, learning the narrowing prediction model using the training dataset, tuning the learned narrowing prediction model using the tuning dataset, and validating the tuned narrowing prediction model using the validation dataset.

[0013] A computer program according to a preferred embodiment of the present invention for achieving the above technical object is stored in a computer-readable storage medium and causes a computer to execute any of the above methods for predicting dialysis access narrowing using a convolutional neural network.

[0014] To achieve the above object, a preferred embodiment of the present invention provides a dialysis access constriction prediction device using a convolutional neural network (CNN) for predicting dialysis access constriction using a convolutional neural network. The device includes: a memory that stores one or more programs for predicting dialysis access constriction using a convolutional neural network (CNN); and one or more processors that perform operations to predict dialysis access constriction using a convolutional neural network (CNN) in accordance with the one or more programs stored in the memory. The processor acquires audio data for the dialysis access of a subject, and predicts the degree of constriction corresponding to the audio data based on a constriction prediction model including a trained convolutional neural network (CNN).

[0015] Here, the processor may preprocess the audio data, input the preprocessed audio data to the narrowing prediction model, and predict the narrowing degree corresponding to the audio data based on the output value of the narrowing prediction model.

[0016] Here, the processor can train the narrowing prediction model based on a training data set including first audio data for the dialysis access route acquired before treatment and second audio data for the dialysis access route acquired after treatment.

[0017] Here, the processor may preprocess the training dataset, set the first audio data as a first correct label, and set the second audio data as a second correct label, and train the narrow prediction model based on the preprocessed training dataset. [Effects of the Invention]

[0018] According to a preferred embodiment of the present invention, a method and apparatus for predicting narrowing of a dialysis access route using a convolutional neural network predicts the degree of narrowing of a dialysis access route from audio data of the subject's dialysis access route based on a narrowing prediction model including a convolutional neural network (CNN), thereby making it possible to more accurately predict the degree of narrowing of the dialysis access route and thereby guide further examinations and treatments.

[0019] The effects of the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. [Brief explanation of the drawings]

[0020] [Figure 1] FIG. 1 is a block diagram illustrating a device for predicting dialysis access narrowing using a convolutional neural network according to a preferred embodiment of the present invention. [Figure 2] 1 is a flowchart illustrating a method for predicting narrowing of a dialysis access route using a convolutional neural network according to a preferred embodiment of the present invention. [Figure 3] 1 is a diagram illustrating a learning process of a narrowing prediction model according to a preferred embodiment of the present invention; [Figure 4] FIG. 4 is a diagram for explaining a preprocessing process of the learning data set shown in FIG. 3. [Figure 5] 1 is a diagram illustrating a process of predicting a degree of constriction using a constriction prediction model according to a preferred embodiment of the present invention; [Figure 6] FIG. 6 is a diagram for explaining a pre-processing process of the audio data shown in FIG. 5. [Figure 7] 10A and 10B are diagrams illustrating an example of a narrowing prediction model learning process and a narrowing degree prediction process according to a preferred embodiment of the present invention; [Figure 8] 1 is a diagram illustrating an example of a spectrogram acquisition process according to a preferred embodiment of the present invention; [Figure 9] FIG. 9 is a diagram showing an example of a spectrogram obtained through the process shown in FIG. 8. [Figure 10] 10A and 10B are diagrams showing examples of spectrograms obtained through the process shown in FIG. 8, where FIG. 10A shows a spectrogram obtained based on audio data for the dialysis access route obtained before treatment, and FIG. 10B shows a spectrogram obtained based on audio data for the dialysis access route obtained after treatment. [Figure 11] 11A and 11B are diagrams illustrating the performance of a narrowing prediction model according to a preferred embodiment of the present invention, in which FIG. 11A shows a confusion matrix and FIG. 11B shows a receiver operation characteristic (ROC) curve. DETAILED DESCRIPTION OF THE INVENTION

[0021] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Advantages and features of the present invention, as well as methods for achieving them, will become apparent from the following detailed description of the embodiments in conjunction with the accompanying drawings. However, the present invention is not limited to the following embodiments, and can be embodied in various forms. These embodiments are provided to fully disclose the present invention and fully convey the scope of the invention to those skilled in the art, and the present invention is defined solely by the claims. Like reference numerals refer to like elements throughout the specification.

[0022] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in the sense that they can be commonly understood by those skilled in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries should not be interpreted ideally or excessively unless specifically defined otherwise.

[0023] In this specification, terms such as "first" and "second" are used to distinguish one component from another, and should not be used to limit the scope of rights. For example, a first component can be named a second component, and similarly, a second component can be named a first component.

[0024] In this specification, the use of identifiers (e.g., a, b, c, etc.) in each step is for convenience of explanation, and the identifiers do not describe the order of each step, and each step may occur in a different order than specified unless the context clearly dictates a specific order. That is, each step may be performed in the same order as specified, may be performed substantially simultaneously, or may be performed in the reverse order.

[0025] As used herein, terms such as "have," "may have," "include," or "may include" refer to the presence of a given feature (e.g., a value, function, operation, or component such as a part) and do not exclude the presence of additional features.

[0026] Hereinafter, preferred embodiments of a method and apparatus for predicting narrowing of a dialysis access route using a convolutional neural network according to the present invention will be described in detail with reference to the accompanying drawings.

[0027] First, with reference to FIG. 1, a device for predicting narrowing of a dialysis access route using a convolutional neural network according to a preferred embodiment of the present invention will be described.

[0028] FIG. 1 is a block diagram illustrating a device for predicting dialysis access narrowing using a convolutional neural network according to a preferred embodiment of the present invention.

[0029] Referring to FIG. 1, a dialysis access route narrowing prediction device 100 using a convolutional neural network according to a preferred embodiment of the present invention can predict the degree of narrowing of a dialysis access route (such as an arteriovenous fistula) of a subject from audio data of the dialysis access route based on a narrowing prediction model including a convolutional neural network (CNN).

[0030] To this end, the narrowing prediction device 100 may include one or more processors 110 , a computer-readable storage medium 130 , and a communication bus 150 .

[0031] The processor 110 can control the operation of the narrowing prediction device 100. For example, the processor 110 can execute one or more programs 131 stored in a computer-readable storage medium 130. The one or more programs 131 may include one or more computer-executable instructions, which, when executed by the processor 110, can be configured to perform operations such that the narrowing prediction device 100 predicts narrowing of a dialysis access route using a convolutional neural network (CNN).

[0032] The computer-readable storage medium 130 is configured to store computer-executable instructions or program code, program data, and / or other suitable forms of information for predicting dialysis access narrowing using a convolutional neural network (CNN). The program 131 stored on the computer-readable storage medium 130 includes a set of instructions executable by the processor 110. In one embodiment, the computer-readable storage medium 130 may be memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, other forms of storage media that can be accessed by the narrowing prediction apparatus 100 and store desired information, or a suitable combination thereof.

[0033] A communication bus 150 interconnects various other components of the narrowing prediction device 100 , including the processor 110 and the computer-readable storage medium 130 .

[0034] The narrowing prediction apparatus 100 may also include one or more input / output interfaces 170 and one or more communication interfaces 190 that provide an interface for one or more input / output devices. The input / output interfaces 170 and the communication interfaces 190 are connected to the communication bus 150. The input / output devices (not shown) can be connected to other components of the narrowing prediction apparatus 100 via the input / output interfaces 170.

[0035] Next, a method for predicting narrowing of a dialysis access route using a convolutional neural network according to a preferred embodiment of the present invention will be described with reference to FIGS.

[0036] FIG. 2 is a flowchart illustrating a method for predicting constriction of a dialysis access route using a convolutional neural network according to a preferred embodiment of the present invention. FIG. 3 is a diagram illustrating a learning process of a constriction prediction model according to a preferred embodiment of the present invention. FIG. 4 is a diagram illustrating a preprocessing process of the learning data set shown in FIG. 3. FIG. 5 is a diagram illustrating a constriction degree prediction process using a constriction prediction model according to a preferred embodiment of the present invention. FIG. 6 is a diagram illustrating a preprocessing process of audio data shown in FIG. 5.

[0037] Referring to FIG. 2, the processor 110 of the narrowing prediction device 100 can learn a narrowing prediction model based on a training dataset (S110).

[0038] Here, the narrowing prediction model includes a convolutional neural network (CNN), which uses a spectrogram as input and outputs a narrowing degree value. For example, the narrowing degree value indicates the probability that the narrowing degree of the dialysis access route is 50% or more, and can have a value between 0 and 1.

[0039] The training data set may include first audio data related to the dialysis access route acquired before angioplasty and second audio data related to the dialysis access route acquired after angioplasty. For example, before angioplasty, audio data (sounds in the audible frequency band between 20 Hz and 1,000 Hz) related to the patient's dialysis access route (such as an arteriovenous fistula) can be acquired using an electronic stethoscope or the like. Similarly, after angioplasty, audio data related to the patient's dialysis access route (such as an arteriovenous fistula) can be acquired using an electronic stethoscope or the like.

[0040] For example, as shown in FIG. 3, the processor 110 can learn the final narrowing prediction model through the steps of "preprocessing process of the learning dataset" → "learning process of the narrowing prediction model" → "tuning process of the narrowing prediction model" → "verification process of the narrowing prediction model."

[0041] That is, the processor 110 may pre-process the training data set.

[0042] Explaining this in more detail with reference to FIG. 4, the processor 110 can pre-process each piece of audio data included in the training data set through the following steps.

[0043] The processor 110 can acquire audio data of a predetermined section from the audio data. For example, the processor 110 can extract audio data of a predetermined section (e.g., 2 seconds to 8 seconds) in order to eliminate the influence of noise, etc.

[0044] The processor 110 may acquire a spectrogram based on audio data of a predetermined section. For example, the processor 110 may convert the audio data into a spectrogram using a Fourier transform (FT).

[0045] The processor 110 may perform normalization on the acquired spectrogram.

[0046] The processor 110 may remove unwanted regions (such as edge boundary regions) from the normalized spectrogram before performing data augmentation.

[0047] The processor 110 may increase the number of normalized spectrograms by horizontally shifting the normalized spectrograms. For example, the processor 110 may increase the number of spectrograms by horizontally shifting the spectrograms multiple times from the time axis.

[0048] The processor 110 may adjust the size of the augmented spectrogram. For example, the processor 110 may adjust the size so that the size of the spectrogram is reduced to a preset size (e.g., 512×512).

[0049] The processor 110 can then train a narrow prediction model based on the pre-processed training dataset, with the first audio data as a first ground truth label and the second audio data as a second ground truth label.

[0050] Here, the first correct answer label indicates that the degree of narrowing of the dialysis access route is 50% or more, and can be set to, for example, "1." The second correct answer label indicates that the degree of narrowing of the dialysis access route is less than 50%, and can be set to, for example, "0."

[0051] More specifically, the processor 110 can learn the narrowing prediction model based on the pre-processed training data set through the following process.

[0052] The processor 110 can divide the pre-processed learning data set into a training data set, a tuning data set, and a validation data set according to a preset criterion, for example, the processor 110 can divide the first audio data set of the learning data set into a training data set, a tuning data set, and a validation data set, and divide the second audio set of the learning data set into a training data set, a tuning data set, and a validation data set according to a preset ratio of 7:2:1.

[0053] The processor 110 may train the narrowing prediction model using a training data set.

[0054] The processor 110 may tune the learned narrowing prediction model using the tuning dataset.

[0055] The processor 110 may validate the tuned narrowing prediction model using a validation dataset.

[0056] The processor 110 can then acquire audio data related to the subject's dialysis access (S130).

[0057] For example, as shown in FIG. 5, the processor 110 can acquire audio data through an "audio data acquisition process" and then an "audio data pre-processing process."

[0058] That is, the processor 110 can acquire audio data related to the dialysis access route of the subject. For example, audio data (sounds in the audible frequency band between 20 Hz and 1,000 Hz) related to the dialysis access route (such as an arteriovenous fistula) of the subject patient for which the degree of narrowing is to be determined can be acquired using an electronic stethoscope or the like.

[0059] The processor 110 can then pre-process the acquired audio data.

[0060] Explaining this in more detail with reference to FIG. 6, the processor 110 can pre-process the audio data through the following steps.

[0061] The processor 110 can acquire audio data of a predetermined section from the audio data. For example, the processor 110 may extract audio data of a predetermined section (e.g., 2 seconds to 8 seconds) in order to eliminate the influence of noise, etc.

[0062] The processor 110 may acquire a spectrogram based on audio data of a predetermined section. For example, the processor 110 may convert the audio data into a spectrogram using a Fourier transform (FT) or the like.

[0063] The processor 110 may perform normalization on the acquired spectrogram.

[0064] The processor 110 may adjust the size of the normalized spectrogram. For example, the processor 110 may adjust the size so that the size of the spectrogram is reduced to a preset size (e.g., 512×512).

[0065] Thereafter, the processor 110 can predict the degree of narrowing corresponding to the audio data based on the trained narrowing prediction model (S150).

[0066] For example, as shown in FIG. 5, the processor 110 can predict the degree of constriction of the subject's dialysis access route through the following steps: "input process of preprocessed audio data" → "process of obtaining output value of constriction prediction model" → "process of predicting degree of constriction."

[0067] That is, the processor 110 may input the pre-processed audio data into a narrow prediction model.

[0068] The processor 110 can then predict the degree of narrowing corresponding to the audio data based on the output value of the narrowing prediction model.

[0069] For example, if the output value of the narrowing prediction model (i.e., narrowing degree value) is "0.95," it indicates that there is a "95%" probability that the narrowing degree of the dialysis access route of the subject is 50% or more. This allows the processor 110 to predict the narrowing degree of the subject as "95%."

[0070] Of course, the processor 110 can compare the output value (i.e., the constriction degree value) of the constriction prediction model with a preset threshold value (e.g., 0.5), and if the output value (i.e., the constriction degree value) is equal to or greater than the threshold value, predict the degree of constriction of the subject as "suspected constriction," or the like, and if the output value (i.e., the constriction degree value) is less than the threshold value, predict the degree of constriction of the subject as "not constricted," or the like.

[0071] Next, an example of a method for predicting narrowing of a dialysis access route using a convolutional neural network according to a preferred embodiment of the present invention and its performance will be described with reference to FIGS.

[0072] FIG. 7 is a diagram illustrating an example of a narrowing prediction model learning process and a narrowing degree prediction process according to a preferred embodiment of the present invention. FIG. 8 is a diagram illustrating an example of a spectrogram acquisition process according to a preferred embodiment of the present invention. FIG. 9 is a diagram illustrating an example of a spectrogram acquired through the process shown in FIG. 8. FIG. 10 is a diagram illustrating an example of a spectrogram acquired through the process shown in FIG. 8. FIG. 10(a) shows a spectrogram acquired based on audio data for a dialysis access route acquired before treatment, and FIG. 10(b) shows a spectrogram acquired based on audio data for a dialysis access route acquired after treatment. FIG. 11 is a diagram illustrating the performance of a narrowing prediction model according to a preferred embodiment of the present invention. FIG. 11(a) shows a confusion matrix, and FIG. 11(b) shows a receiver operation characteristic (ROC) curve.

[0073] Referring to FIG. 7, an example of a method for predicting constriction of a dialysis access route using a convolutional neural network according to a preferred embodiment of the present invention may broadly include a constriction prediction model learning process consisting of an "image preprocessing process (Image Preprocessing shown in FIG. 7)" and a "deep learning process (Deep Learning Process shown in FIG. 7)," and a constriction degree prediction process consisting of a "patient audio data preprocessing process (User shown in FIG. 7)," a "deep learning process (Deep Learning Process shown in FIG. 7)," and a "patient constriction degree prediction process (Output shown in FIG. 7)."

[0074] The data input to the narrowing prediction model (audio data included in the training dataset used to train the narrowing prediction model or audio data of the object for which the degree of narrowing is to be determined) may be a mel spectrogram, a 512x512 image file with three RGB channels.

[0075] It is possible to obtain audio files that record the patient's dialysis access route (arteriovenous fistula, etc.) being listened to using an electronic stethoscope, etc. The recordings were made for approximately 10 seconds, but when a person records directly, the playback time of each audio file can vary, and noise such as touching the stethoscope when starting and stopping recording can be introduced into the audio file. To remove such noise, we actually used the time from 2 to 8 seconds of each audio file, i.e., 6 seconds of audio data.

[0076] [Table 1]

[0077] The audio file can then be sampled according to a particular sampling rate and the numeric numbers can be stored in the form of an array.

[0078] [Table 2]

[0079] Here, sr can be set to a specific value, but since the present invention uses the native sampling rate, sr = Non. This allows for the creation of a graph of sampling interval (x-axis, time) vs. amplitude (y-axis), but this is not very useful for analysis. Sound can basically be considered as the sum of sine functions with specific frequencies. Frequency analysis of the previously obtained y waveform can reveal how each frequency component is structured at a specific time. This method is called the Fourier transform (FT). By solving the Fourier equation, an amplitude vs. time graph can be converted into a frequency vs. time graph. In this invention, the STFT (short time Fourier transform) is used as the Fourier transform (FT).

[0080] [Table 3]

[0081] Here, fmax is the maximum frequency that determines the analysis range, and according to Nyquist's law, the maximum frequency is usually determined by the value of sampling rate / 2. When performing STFT in this way, the number of divisions for analysis can be determined by the hop length shown in Figure 8. n_fft is the FFT length (or window length) to be analyzed, which was determined to be 25 msec. The hop length was set to 10 msec, with 15 msec (overlap length) overlapping per square.

[0082] Using this method, a mel spectrogram can be obtained, as shown in Figure 9. The X axis is time, the Y axis is frequency, and the intensity (decibels) of a specific frequency in a specific time period can be represented by color.

[0083] However, when training using an actual mel spectrogram, feature extraction is required, and spectrogram normalization is required to improve the ability to distinguish spectrogram power, which is represented by color in the mel spectrogram, and to ensure uniformity of the data. In other words, spectrogram normalization is necessary because the amount of noise can vary from recording to recording, and sound waves with greater power can be recorded at certain frequencies depending on the degree of constriction, and to ensure that features can be recognized as well as possible when training a constriction prediction model.

[0084] [Table 4]

[0085] Using the above method, a mel spectrogram can be obtained from each audio file. Figure 10 shows examples of spectrograms before and after angioplasty. Figure 10(a) is the mel spectrogram before angioplasty, and Figure 10(b) is the mel spectrogram after angioplasty. After the procedure, the degree of narrowing of the dialysis access route (such as an arteriovenous fistula) improves, and a spectrogram with greater power at higher frequencies can be seen. In fact, listening to the audio recordings before and after the procedure confirms that the audio after the procedure is louder and more audible. The mel spectrogram obtained in this way was labeled with the first correct answer label "pre(1)" and saved in a folder if it was obtained before the angioplasty procedure, because the degree of narrowing was 50% or more (calculated by comparing the diameter values ​​of the narrowed area of ​​the arteriovenous fistula with that of a normal blood vessel in actual angiography). On the other hand, the mel spectrogram obtained after the angioplasty procedure, because the degree of narrowing was less than 50% (when the degree of narrowing was confirmed to be less than 50% in actual angiography), was labeled with the second correct answer label "post(0)" and saved in a folder.

[0086] And there is a white boundary surrounding the edge region of the mel spectrogram, as shown in Figure 9. To make this boundary more clearly visible, the blue border line has been arbitrarily added.

[0087] To train a narrowing prediction model, the amount of mel spectrogram data must be increased, using the horizontal shifting method. When developing a convolutional neural network for cat recognition, it is possible to amplify the original cat photo by using techniques such as vertical / horizontal flip or by applying various angles to the original photo. However, unlike typical cat photos, mel spectrograms are vectorgrams with fixed values ​​and meanings on the x-, y-, and z-axes. Therefore, horizontal shifting is the only method that can increase the amount of data. Data obtained through horizontal shifting can be used as training data because it is the result of recording the same patient at different start and end times. In other words, the reason for using horizontal shifting is that in reality, mel spectrograms can shift along the x-axis (time) depending on the start and end times of the recording. If we consider the recurring peaks seen in a mel spectrogram to be sin(x) or cos(x) functions, then depending on how we set the recording time range, the captured wave may look like sin(x+a) or cos(x+a).

[0088] To increase the data volume, we used the following ImageDataGenerator.

[0089] [Table 5]

[0090] In the code above, we can see that the width shift range is set to 0.9 for horizontal shifting. In the code above, the increment is set to 50 (i>50), meaning that one mel spectrogram image is taken and moved along the x-axis to generate 50 images, but 50 is an arbitrary value and does not have to be limited to 50. However, as shown in Figure 9, if there are white edges, the white vertical lines on the left and right edges may appear in the middle when horizontal shifting, damaging the mel spectrogram data. Therefore, we perform preprocessing to remove these white edges.

[0091] [Table 6]

[0092] The invisible black lines at the left and right edges of the white border area are also removed.

[0093] [Table 7]

[0094] The mel spectrogram obtained in this way is confirmed to have a size of 2328 x 909 using the following method.

[0095] [Table 8]

[0096] If the image size is too large, when the narrowing prediction model compresses the array in stages, the mel spectrogram values ​​before and after processing may not show a significant difference when a large area from the starting image is compressed and the final layer is reached. Also, the learning time of the narrowing prediction model will be longer, and the rectangular image will be resized to a square photograph.

[0097] [Table 9]

[0098] The above code is an example of adjusting the size of the mel spectrogram to 512x512, and after adjusting the size as described above, it will be input into the narrowing prediction model.

[0099] The training dataset was split into training, tuning, and validation datasets in a ratio of 7:1:2. For this purpose, we used the train_test_split function, which randomly splits an array containing a list of files in a folder into training, tuning, and validation subsets.

[0100] [Table 10]

[0101] Then, for example, filelist_tune will contain the following melspectrogram.png file in the form of an array:

[0102] [Table 11]

[0103] If the above file is set to pre=0 and post=1, it will become a binary array of the following form:

[0104] [Table 12]

[0105] In other words, in the case of aug_0_4960.png in the file list of tuning, it is stored as 0 because it is pre (audio recording taken before treatment), and in the case of aug_0_1335.png, it is stored as 1 because it is post (audio recording taken after treatment).The performance of the narrowing prediction model according to the present invention was tested using the ResNET50 model, a convolutional neural network (CNN) model.

[0106] ResNET50, like a typical convolutional neural network (CNN) model, consists of an input layer, convolutional layer, max pooling layer, average pooling layer, and output layer. Here, the convolutional layer consists of 50 layers and extracts image features from the mel spectrogram. The max pooling layer sub-samples the features extracted from the convolutional layer to improve the stability and efficiency of the system. The average pooling layer reduces the number of parameters. The output layer outputs the following values:

[0107] That is, the values ​​output through the output layer can be values ​​for the predictive ability and diagnostic performance of the narrowing prediction model for dialysis access narrowing of 50% or more. For example, as shown in the following example, values ​​for sensitivity, specificity, positive predictive value, negative predictive value, accuracy, etc. can be output. Based on this, a confusion matrix and receiver operation characteristic (ROC) curve can be obtained as shown in FIG. 11, and the area under the curve (AUC) value of the diagnostic ability can be calculated through the ROC curve.

[0108] [Table 13]

[0109] The mel spectrograms of specific patients included in the validation dataset can be used to obtain a YES / NO result on whether or not a 50% or greater dialysis access route (such as an arteriovenous fistula) narrowing should be suspected. When the narrowing prediction model is run, the output shows whether each mel spectrogram is narrowed by 0 or 1, indicating whether the narrowing is less than 50% or greater than 50%.

[0110] [Table 14]

[0111] In the case of the ResNet50 model, the network trains to minimize H(x)-x so that the output value is x, resulting in an output of "0.94346315". This is a value close to 1, and in such cases, it can be considered that there is a 50% or greater narrowing. In such cases, print("YES") can be output when starting the model. Conversely, in the following case, the value is close to 0, so it can be recognized as a narrowing of less than 50%, rather than 0 or a 50% or greater narrowing, and print("NO") can be output.

[0112] [Table 15]

[0113] If significant narrowing of 50% or more is suspected, additional testing for narrowing of the dialysis access route (such as arteriovenous fistula) is recommended, and if the answer is "YES," a recommendation such as "Severe narrowing of the hemodialysis access route is suspected, so hemodialysis may not be performed properly. Additional tests such as Doppler ultrasound or angiography are required, so please visit a nearby hospital" can also be output. If the mel spectrogram of a specific patient included in the validation dataset is suspected to have a narrowing of the dialysis access route (such as arteriovenous fistula) of 50% or more, the degree of suspicion of the narrowing prediction model can be output as a percentage.

[0114] [Table 16]

[0115] For this patient's mel spectrogram, the narrowing prediction model is about 94%, which means that it predicts that there is a narrowing of more than 50%. The training process, tuning process, and validation process of the narrowing prediction model using the ResNet50 model shown in the code below are as follows:

[0116] [Table 17]

[0117] Batch_size is the number of samples used for training, and epoch is the number of times the training goes back and forth between the 50 layers of the ResNet. In other words, epochs=10 means that training will be done using the base data 10 times. These values ​​are not fixed, and batch_size and especially the epoch value must be adjusted several times to optimize the model. For epoch, if the value is too small, the model will tend to underfit the data, while if it is too large, it will overfit. For example, if there are 100 mel spectrograms, the batch size is 20, so 20 pieces of data will be trained per iteration, so 1 epoch = 100 / batch size = 5 iterations. 40 epochs will require 200 iterations.

[0118] The optimizer used to update the model after each learning step was Keras SGD (stochastic gradient descent). There are many types of optimizers available, including RMSprop, Adam, and Adadelta. While SGD was used in this invention, it is not limited to SGD. When training a narrow prediction model, the learning rate is generally set to a value between 0.1 and 0.01, and momentum is often set to 0.9. In this invention, the learning rate was set to 0.02.

[0119] [Table 18]

[0120] The basic principle of optimizing is to make the learning rate "large at first, and then smaller and smaller" (Reference: Qian Ning, On the momentum term in gradient descent learning algorithms, Neural networks 12.1(1999):145-151). Momentum has the same value of the learning rate itself, but when changing parameters, it uses an adjustment term called the momentum term to similarly express the concept of "large at first, and then smaller and smaller."

[0121] The parameter of the neural network model for the error function E is θ, and the slope of E with respect to θ is ∇θ E , parameter difference Δθ (t) If we use equation (1) to express the parameter change using momentum at step t, then equation (2) is given. γΔθ (t-1) :Momentum term The coefficient γ (<1) is generally set to a value such as 0.5 or 0.9. Δθ (t) =Δθ (t) -γΔθ (t-1) (1) Δθ (t) =-η∇ θ E(θ)+γΔθ (t-1) (2)

[0122] In other words, by setting the learning rate and momentum in this way, accuracy can be improved quickly in the initial epochs.

[0123] After setting all the parameters in this way, we fit or train the narrowing prediction model as follows:

[0124] [Table 19]

[0125] Since the mel spectrogram has already been separated and saved into pre and post files, the model is created while learning, with 1 when reading from PRE_PATH+filename and 0 when reading from POST_PATH+filename. During the tuning or fine-tuning stage, the model with the highest accuracy for the tuning-set data is selected based on the accuracy observed when the completed model for each epoch is input with the tuning-set data. For example, if you run the program with epochs=10 for testing, you will get the following results.

[0126] [Table 20]

[0127] Here, accuracy is the accuracy of the model for the training set, and val_accuracy is the accuracy of the model for the tuning set. As explained above, a small epoch value can lead to underfitting, while a large epoch value can lead to overfitting. Therefore, from epoch 1 / 10 to epoch 10 / 10, accuracy improves significantly (0.5981 → 0.9471). However, val_accuracy for the tuning set peaks at epoch 9 / 10 and then decreases slightly to 0.800 at epoch 10 / 10. This is because the training set melspectrogram was overfitted, resulting in poor fitting and a decrease in accuracy when the tuning set melspectrogram was input. Therefore, the tuning stage typically involves determining the model for the epoch with the best accuracy and val_accuracy. In the example above, tuning involves determining the epoch 9 / 10 model. So, after determining the Epoch 9 train weights, we apply the model to the validation set to see how accurately it makes predictions.

[0128] [Table 21]

[0129] Then, the output value described above can be obtained.

[0130] The operations according to the present embodiment may be embodied in the form of program instructions executable by various computer means and recorded on a computer-readable storage medium. A computer-readable storage medium refers to any medium involved in providing instructions to a processor for execution. The computer-readable storage medium may include program instructions, data files, data structures, or combinations thereof. Examples include magnetic media, optical recording media, and memories. The computer program may be distributed over computer systems connected via a network, and the computer-readable code may be stored and executed in a distributed manner. Functional programs, codes, and code segments for implementing the present embodiment may be easily construed by programmers skilled in the art to which the present embodiment pertains.

[0131] The present embodiment is intended to explain the technical idea of ​​the present embodiment, and does not limit the scope of the technical idea of ​​the present embodiment. The scope of protection of the present embodiment should be interpreted according to the following claims, and all technical ideas within the scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment. [Explanation of symbols]

[0132] 100...Narrowing prediction device 110 Processor 130 Computer-readable storage medium 131 Program 150 Communication Bus 170 Input / Output Interface 190 Communication Interface

Claims

1. A method for predicting stenosis of a dialysis access route using a dialysis access route stenosis prediction device, comprising: acquiring audio data relating to a dialysis access route of a subject; predicting a degree of narrowing corresponding to the audio data based on a narrowing prediction model including a trained convolutional neural network (CNN) by the narrowing prediction device; A method for predicting narrowing of a dialysis access route using a convolutional neural network, further comprising a step of training the narrowing prediction model based on a training dataset including first audio data for the dialysis access route acquired before treatment and second audio data for the dialysis access route acquired after treatment.

2. The step of acquiring audio data includes: preprocessing the audio data; The step of predicting the degree of narrowing includes: The method further comprises inputting the pre-processed audio data into the narrowing prediction model by the narrowing prediction device, and predicting a narrowing degree corresponding to the audio data based on an output value of the narrowing prediction model. A method for predicting narrowing of a dialysis access route using the convolutional neural network of claim 1.

3. The step of acquiring audio data includes: obtaining the audio data of a predetermined section from the audio data, obtaining a spectrogram based on the audio data of the predetermined section, normalizing the obtained spectrogram, and adjusting the size of the normalized spectrogram. A method for predicting narrowing of a dialysis access route using the convolutional neural network of claim 2.

4. The narrowing prediction model is A spectrogram is input and a narrowing degree value is output. A method for predicting narrowing of a dialysis access route using the convolutional neural network of claim 1.

5. The step of training the narrowing prediction model includes: preprocessing the training dataset; The first audio data is designated as a first correct label, the second audio data is designated as a second correct label, and the narrow prediction model is trained based on the pre-processed training dataset.

5. The method for predicting narrowing of a dialysis access route using a convolutional neural network according to claim 4, comprising:

6. The step of training the narrowing prediction model includes: For each of the audio data included in the training dataset, acquiring the audio data of a predetermined section from the audio data, acquiring a spectrogram based on the audio data of the predetermined section, normalizing the acquired spectrogram, horizontally shifting the normalized spectrogram to increase the number of spectrograms, and adjusting the size of the increased spectrograms to preprocess the training dataset. A method for predicting narrowing of a dialysis access route using the convolutional neural network of claim 5.

7. The step of training the narrowing prediction model includes: Dividing the pre-processed learning data set into a training data set, a tuning data set, and a validation data set according to a predetermined criterion; training the narrowing prediction model using the training dataset, tuning the trained narrowing prediction model using the tuning dataset, and validating the tuned narrowing prediction model using the validation dataset. A method for predicting narrowing of a dialysis access route using the convolutional neural network of claim 5.

8. A computer program stored in a computer-readable storage medium for causing a computer to execute the method for predicting dialysis access narrowing using a convolutional neural network according to claim 1.

9. A narrowing prediction device for predicting narrowing of a dialysis access route using a convolutional neural network (CNN), comprising: a memory storing one or more programs for predicting dialysis access narrowing using a convolutional neural network (CNN); one or more processors that perform operations to predict dialysis access narrowing using a convolutional neural network (CNN) in accordance with the one or more programs stored in the memory; Including, The processor: obtaining audio data relating to a dialysis access route for the subject; predicting a degree of narrowing corresponding to the audio data based on a narrowing prediction model including a trained convolutional neural network (CNN); training the narrowing prediction model based on a training data set including first audio data for the dialysis access route acquired before the procedure and second audio data for the dialysis access route acquired after the procedure; A device for predicting dialysis access narrowing using a convolutional neural network.

10. The processor: preprocessing the audio data; inputting the pre-processed audio data into the narrowing prediction model, and predicting a narrowing degree corresponding to the audio data based on an output value of the narrowing prediction model; A device for predicting narrowing of a dialysis access route using the convolutional neural network according to claim 9.

11. The processor: preprocessing the training dataset; The first audio data is designated as a first correct label, and the second audio data is designated as a second correct label, and the narrow prediction model is trained based on the pre-processed training dataset. A device for predicting narrowing of a dialysis access route using the convolutional neural network according to claim 9.

Citation Information

Patent Citations

  • Audio recognition method and system

    JP2018534609A

  • JPP6712028B

  • JPP6854554B

  • Shunt murmur analysis device, shunt murmur analysis method, computer program, and recording medium

    WO2016207951A1

  • Managing respiratory conditions based on sounds of the respiratory system

    WO2019229543A1