Biological information analysis system and biological information analysis program
Patent Information
- Application Number
- PCT/JP2026/006760
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-25
- Filing Date
- 2026-02-25
- Publication Date
- 2026-09-03
Smart Images

Figure JP2026006760_03092026_PF_FP_ABST
Abstract
Description
Biomedical information analysis system and biomedical information analysis program
[0001] This invention relates to a series of biological information analysis systems and biological information analysis programs that classify data based on biological information obtained from a subject.
[0002] The heart repeatedly contracts and expands, acting as a pump to send blood throughout the body. This activity (hereinafter referred to as cardiac activity) is maintained by weak electrical impulses in myocardial cells, but if abnormalities appear in these electrical impulses, abnormalities occur in the heart's activity as well. This is called arrhythmia. There are various types of arrhythmias, some of which are highly lethal, causing cardiac arrest, and others that can lead to serious diseases such as stroke. Thus, the importance of arrhythmias differs depending on the type and frequency. Therefore, detecting the occurrence of arrhythmias and classifying them is extremely important from a healthcare perspective. To diagnose these arrhythmias, techniques are known that use biological information obtained from patients by biological information measuring devices (see, for example, Patent Document 1).
[0003] Figures 21A to 21C show examples of electrocardiogram (ECG) signals. The ECG waveform 100 shown in Figure 21A is a normal ECG waveform without arrhythmia. The ECG waveform 110 shown in Figure 21B shows an ECG waveform with premature atrial contraction (PAC). The ECG waveform 120 shown in Figure 21C shows an ECG waveform with premature ventricular contraction (PVC). The ECG waveform may contain waveform components called P wave 101, Q wave 102, R wave 103, S wave 104, and T wave 105 (see Figure 21A). Each of these waveforms has characteristic features in terms of the interval between occurrences and waveform shape, which allows for the diagnosis of arrhythmia.
[0004] In recent years, technologies have been developed that use biometric information analysis devices to automatically analyze biological data and diagnose arrhythmias. This technology is expected to reduce the burden on doctors, prevent missed diagnoses of arrhythmias by handling longer-term biological data, and lead to highly accurate diagnoses.
[0005] Japanese Unexamined Patent Application Publication No. 2023-068219
[0006] By the way, as in Patent Document 1, calculating feature quantities such as the occurrence intervals and waveform shapes of electrocardiographic waveforms and analyzing them is effective for detecting arrhythmia and classifying arrhythmia types, but the following problems are raised. It is known that electrocardiographic waveforms have characteristics unique to individual patients (individuality), and that even in the same patient, the shape and rhythm vary depending on the time zone and condition (e.g., during exercise, at rest, stress, etc.) (variability). In Patent Document 1, for example, feature quantities related to the waveform shapes of QRS waves and P waves are used. In order to perform high-precision classification in consideration of such individuality and variability, multidimensional analysis using a larger number of feature quantities may be required. Furthermore, regarding performing multidimensional analysis that takes into account individuality and variability, it is important to adjust the contribution of each feature quantity according to the data group, and further to have high adaptability to individuality and variability, but Patent Document 1 does not take this point into consideration. As described above, in the technique of classifying waveforms based on feature quantities of electrocardiographic waveforms as in Patent Document 1, there has been a need to improve classification accuracy.
[0007] An object of the present invention is to solve the above problems and provide a biological information analysis system and a biological information analysis program capable of classifying waveforms represented by biological information with high accuracy.
[0008] The biological information analysis system of the present invention that achieves the above object has the following configuration. (1) A biological information analysis system that classifies data using target biological information, comprising: an input reception unit that receives input of the biological information; a waveform detection unit that detects waveform data from the biological information; a feature quantity acquisition unit that acquires a plurality of feature quantities related to the waveform data; a dimension compression unit that extracts features from the feature quantities using a dimension compression model; a clustering unit that clusters the feature extraction results using a clustering model; a grouping unit that groups the waveform data based on the clustering results; A biological information analysis system comprising:
[0009] (2) The biological information analysis system according to (1), wherein the biological information is information indicating the activity of the target heart, and the biological information is used to classify arrhythmias.
[0010] (3) A screening unit for extracting waveform data from the waveform data to be classified, further comprising the biological information analysis system according to (1) or (2).
[0011] (4) The dimensionality reduction model is a nonlinear dimensionality reduction model, the bio-information analysis system described in any one of (1) to (3).
[0012] (5) The clustering model is a nonlinear clustering model, the bio-information analysis system described in any one of (1) to (4).
[0013] (6) The biomedical information analysis system according to any one of (1) to (5), wherein the feature quantities include at least one of the following: time-frequency representation data obtained by applying a time-frequency transformation to the waveform data; shape data obtained from the width of the waveform components in the waveform data; and periodic data relating to the interval between preceding and succeeding waveform data in the waveform data.
[0014] (7) A biological information analysis system according to any one of (1) to (6), further comprising: an extraction unit that extracts waveform data belonging to a specified group based on the grouping results; and a control unit that performs featureization by the dimensionality compression unit, clustering by the clustering unit, and grouping by the grouping unit on the waveform data belonging to the specified group.
[0015] (8) The dimensionality reduction model is characterized in that the feature quantities are characterized in two dimensions, and the characteristic result is two-dimensional numerical data, as described in any one of (1) to (7).
[0016] (9) The dimensionality reduction model is characterized in that the feature quantities are characterized in three dimensions, and the characteristic result is three-dimensional numerical data, as described in any one of (1) to (7).
[0017] (10) A biological information analysis system according to any one of (1) to (9), further comprising a labeling unit that labels each group with a waveform type based on the grouping results.
[0018] (11) The biological information analysis system according to (10), wherein the labeling unit selects a representative waveform data from the waveform data belonging to each group for each group, determines the waveform type of the representative waveform data, and labels all the waveform data of the group to which the representative waveform data belongs with the waveform type.
[0019] (12) The biological information analysis system according to (10), wherein the labeling unit determines a waveform type based on a statistical amount calculated for each group and labels all waveform data belonging to each group with the waveform type.
[0020] (13) The biological information analysis system according to (10), further comprising a labeled data input unit that accepts labeled waveform data, wherein the labeling unit characterizes the feature quantities obtained from the labeled waveform as labeled waveform data using the dimensionality reduction model to obtain a labeled featureization result, and labels the waveform type for each group based on the labeled featureization result.
[0021] (14) The biological information analysis system according to (1), wherein the dimensionality reduction unit is composed of a first dimensionality reduction unit and a second dimensionality reduction unit, the first dimensionality reduction unit obtains a first featureization result by characterizing a part of the feature quantities using a dimensionality reduction model, the second dimensionality reduction unit obtains a second featureization result by characterizing any of the feature quantities using a dimensionality reduction model, including at least the first featureization result, and the clustering unit clusters the second featureization result using a clustering model.
[0022] (15) A bio-information analysis program that classifies data using target bio-information, wherein the program causes a computer to perform a series of processes, including: receiving input of bio-information; detecting waveform data from the bio-information; obtaining a plurality of feature quantities relating to the waveform data; characterizing the feature quantities using a dimensionality reduction model; clustering the featureization results using a clustering model; and grouping the waveform data based on the clustering results.
[0023] According to the present invention, waveform data obtained from biological information can be characterized using a dimensionality reduction model, and similar waveforms can be grouped together to classify waveform types, such as arrhythmias, with high accuracy.
[0024] Figure 1 is a diagram showing an example configuration of a bio-information analysis system according to Embodiment 1. Figure 2 is a flowchart showing the waveform type classification process according to Embodiment 1. Figure 3 is a flowchart showing the feature acquisition process by the feature acquisition unit according to Embodiment 1. Figure 4 is a diagram illustrating an example of time-frequency representation data (scalogram) obtained by applying a continuous wavelet transform to an electrocardiogram waveform. Figure 5 is a diagram illustrating an example of a featureization result obtained by three-dimensionally characterizing the acquired features using UMAP in the dimensionality reduction unit. Figure 6 is a diagram illustrating an example of a clustering result obtained by clustering the featureization results using DBSCAN in the clustering unit. Figure 7 is a diagram showing representative electrocardiogram waveforms in each cluster clustered by the clustering unit according to Embodiment 1. Figure 8 is a diagram illustrating an example of a grouping result output by the grouping unit according to Embodiment 1. Figure 9 is a diagram showing an example configuration of a bio-information analysis system A according to Embodiment 2. Figure 10 is a flowchart showing the waveform type classification process according to Embodiment 2. Figure 11 is a diagram illustrating an example of the featureization results for each labeled electrocardiogram waveform used for labeling each group in the labeling unit according to Embodiment 2. Figure 12A is an enlarged view (part 1) of a part of the featureization results shown in Figure 11. Figure 12B is an enlarged view (part 2) of a part of the featureization results shown in Figure 11. Figure 12C is an enlarged view (part 3) of a part of the featureization results shown in Figure 11. Figure 13 is a diagram illustrating an example of the configuration of the biological information analysis system according to Embodiment 3. Figure 14 is a flowchart illustrating the waveform type classification process according to Embodiment 3. Figure 15 is a diagram illustrating an example of the configuration of the biological information analysis system according to Embodiment 4. Figure 16 is a flowchart illustrating the waveform type classification process according to Embodiment 4. Figure 17 is a diagram illustrating an example of the configuration of the biological information analysis system D according to Embodiment 5. Figure 18 is a flowchart illustrating the waveform type classification process according to Embodiment 5. Figure 19 is a diagram illustrating an example of the clustering results obtained by the clustering of the additional analysis by the clustering unit according to Embodiment 5.Figure 20 shows representative electrocardiogram waveforms for each of the three clusters clustered by the clustering of the additional analysis performed by the clustering unit according to Embodiment 5. Figure 21A is a diagram showing an example of an electrocardiogram signal (part 1). Figure 21B is a diagram showing an example of an electrocardiogram signal (part 2). Figure 21C is a diagram showing an example of an electrocardiogram signal (part 3).
[0025] Embodiments of the biological information analysis system according to the present invention will be described in detail below with reference to the drawings. However, the present invention is not limited by these embodiments. Furthermore, the individual embodiments of the present invention are not independent but can be combined and implemented as appropriate.
[0026] (Embodiment 1) Figure 1 is a diagram showing an example of the configuration of a biological information analysis system according to Embodiment 1 of the present invention. The biological information analysis system comprises an electrocardiogram signal measuring device 2, which is a biological information measuring device, an analysis system 3, a receiving terminal 4, and a learning device 5.
[0027] The biological information measuring device is freely determined and not limited to the biological information to be acquired. Furthermore, it is assumed that the biological information measuring device is equipped with a connector for electrical connection to the analysis system 3, a communication device equipped with means for communication with the biological information analysis system, or a biological information output mechanism such as an input port for a medium on which the acquired biological information is written. In this embodiment 1, the biological information is an electrocardiogram signal 10, the biological information measuring device is an electrocardiogram signal measuring device 2, the electrocardiogram signal measuring device 2 is equipped with an input port for a medium on which the acquired biological information, the electrocardiogram signal 10, is written, and the subject 1 is a person as an example.In this embodiment 1, an example of acquiring an electrocardiogram signal 10 and classifying the electrocardiogram shape of the subject 1 is described, but other than the electrocardiogram signal 10, it is not particularly limited as long as it is possible to estimate the heart rate cycle, and may be a pulse wave signal, heart sounds, etc.
[0028] The electrocardiogram signal 10 to be analyzed is acquired from the subject 1 via the electrocardiogram signal measuring device 2. The subject 1 is not particularly limited to a person or an animal. The receiving terminal 4 is, for example, a smartphone held by the subject 1 or a caregiver (for example, a surgeon such as a doctor, a factory hygiene manager, or a nearby worker).
[0029] An example of an electrocardiogram signal measuring device 2 is a wearable electrocardiograph. Specifically, the electrocardiogram signal measuring device 2 comprises a garment body worn on a subject 1, multiple electrodes, and an electrocardiograph 100 that is electrically connected to each electrode. The electrocardiograph 100 and the like are fixed to the garment body using a band or the like.
[0030] The electrocardiograph 100 is an example of a measuring device that acquires electrocardiogram signals 10. The electrocardiograph 100 has the function of continuously acquiring the electrocardiogram signals 10 of a subject 1, the function of storing the acquired electrocardiogram signals 10, and the function of transferring data to the analysis system 3 by communication. Alternatively, the electrocardiograph 100 may transfer data including the electrocardiogram signals 10 to a server device, and the server device may transfer the electrocardiogram signals 10 to the analysis system 3. Hereinafter, subject 1 will be described as a person wearing the electrocardiogram signal measuring device 2.
[0031] In this embodiment 1, the electrocardiogram signal 10 acquired from the subject 1 includes information about the subject 1, the date and time of acquisition, the location, and the acquisition result. The acquisition result is waveform data 11 obtained by plotting time on the horizontal axis and voltage on the vertical axis (see Figure 1). The waveform data 11 may include waveform components called P waves, Q waves, R waves, S waves, and T waves, and their shapes are used as criteria for determining arrhythmias, etc. For example, the waveform data 11 is made up of repeating patterns of these waveform components. Note that the shape of the waveform data 11 contained in the electrocardiogram signal 10 may change depending on the subject, their condition, and how the electrocardiograph 100 is attached.
[0032] Returning to Figure 1, the analysis system 3 comprises a communication unit 311, a waveform detection unit 312, a feature acquisition unit 313, a dimensionality compression unit 314, a clustering unit 315, a grouping unit 316, an input / output unit 317, a control unit 318, and a storage unit 319.
[0033] The communication unit 311 can communicate with the electrocardiogram signal measuring device 2 and the receiving terminal 4 via a communication network. The communication network referred to here is configured using, for example, an existing public telephone network, LAN (Local Area Network), WAN (Wide Area Network), etc., and can be wired or wireless. The communication unit 311 is composed of, for example, a connector that electrically connects to the communication target, a communication device equipped with means for communicating with the electrocardiogram signal measuring device 2, or an input port for a media on which data is stored. The communication unit 311 functions as an input receiving unit that accepts input of biological information.
[0034] The waveform detection unit 312 detects waveform data from the input biological information. In this embodiment 1, since the biological information is an electrocardiogram signal, the waveform data here refers to the electrocardiogram waveform. The detection process by the waveform detection unit 312 may include methods such as detecting characteristic peaks in the electrocardiogram waveform, such as the R wave, by differential processing or thresholding, or pattern recognition using template matching or machine learning models, but is not limited to these, and may also include preprocessing such as bandpass filtering.
[0035] The feature acquisition unit 313 acquires multiple features for each detected waveform data. The features are not particularly limited, but it is desirable that they represent the waveform type, such as time-series information like the time of occurrence of the waveform data. Examples include, but are not limited to, multidimensional data obtained by applying arbitrary processing to the waveform data, shape data obtained from each waveform component of the waveform data, and periodic data relating to the interval between preceding and succeeding waveform data in the waveform data. Specifically, multidimensional data includes time-frequency representation data obtained by applying time-frequency transformation to the waveform data, and recurrence plot data. Specifically, shape data includes the waveform width, wave height, occurrence time, and interval between occurrence times for each waveform component (P wave, Q wave, R wave, S wave, and T wave) of the waveform data. Specifically, periodic data includes the interval between the target waveform data and the waveform data immediately preceding it, the interval with the waveform data immediately following it, and their ratios. The feature acquisition processing by the feature acquisition unit 313 is not limited to the processing method depending on the type of features to be acquired.
[0036] Here, it is desirable to apply preprocessing to the features according to the specifications of the dimensionality reduction model used in the dimensionality reduction unit 314. More specifically, it is preferable to convert multidimensional data into one-dimensional vectors suitable for input to the dimensionality reduction model using any vectorization means. Examples of vectorization means include, but are not limited to, smoothing to one dimension or extracting feature vectors from the intermediate layer using any trained model. Furthermore, it is desirable that the total number of dimensions in the features is greater than the number of dimensions characterized by the dimensionality reduction unit 314 described later. In this embodiment 1, for example, feature vectors extracted by a trained model are used for time-frequency representation data obtained by applying time-frequency transformation to waveform data, but are not limited to this.
[0037] The dimensionality reduction unit 314 characterizes the acquired features to a specified number of dimensions using a dimensionality reduction model and outputs the characterization results. The dimensionality reduction model can be a linear dimensionality reduction model such as PCA (Principal Component Analysis), or a nonlinear dimensionality reduction model such as UMAP (Uniform Manifold Approximation and Projection) or t-SNE (t-distributed Stochastic Neighbor Embedding), but is not limited to these. The number of dimensions to be characterized can be set according to the purpose of the analysis, for example, 2 or 3, but is not limited to these. In this embodiment 1, the explanation will be given using UMAP, one of the nonlinear dimensionality reduction models, to characterize in three dimensions, but is not limited to this.
[0038] The clustering unit 315 clusters the featureization results using a clustering model and outputs the clustering results. Here, the clustering model can be a linear clustering model such as K-means or GMM (Gaussian Mixture Model), or a nonlinear clustering model such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise) or HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise). In this embodiment 1, an example of clustering using DBSCAN, one of the nonlinear clustering models, will be described, but the embodiment is not limited to this.
[0039] The grouping unit 316 groups the waveform data based on the clustering results and outputs the grouping results. The grouping performed here refers to assigning a group number to each waveform based on the clustering results output by the clustering unit 315. The processing method and output format of the grouping results by the grouping unit 316 are determined according to operational requirements and are not limited to these.
[0040] The input / output unit 317 is composed of devices having input and output functions, and outputs various information under the control of the control unit 318. The input / output unit 317 includes a user interface such as a keyboard, mouse, and microphone, and has an input function that accepts the electrocardiogram signal 10 output from the electrocardiogram signal measuring device 2. Furthermore, the input / output unit 317 has an output function such as a display made of liquid crystal or organic EL (Electro Luminescence), or a speaker that outputs sound.
[0041] Furthermore, it is assumed that the input / output unit 317 will have an input mechanism for biometric information, such as an input port for a media on which biometric information is written. If the input function of the input / output unit 317 overlaps with the function of the communication unit 311, it is sufficient for either unit to have that function.
[0042] Furthermore, the input method for the biological information input to the analysis system 3 can be selected according to operational requirements. In addition to using previously acquired electrocardiogram signals 10, the acquisition of electrocardiogram signals 10 from the electrocardiogram signal measuring device 2 and input to the analysis system 3 may be performed in parallel. In this case, the electrocardiogram signal measuring device 2 may continuously transmit values acquired in a continuous manner, or batch values acquired in a batch manner may be transmitted in batches at regular intervals, and is not particularly limited.
[0043] The control unit 318 generally controls the operation of the analysis system 3. The control unit 318 also transmits the grouping result obtained by the grouping unit 316, together with wearer information, to the receiving terminal 4 via the communication unit 311. The wearer information is information for identifying the subject 1, including personal information such as the name and gender of the subject 1, and the management number of the wearable device or peripheral device worn by the subject 1. Furthermore, the control unit 318 may output the grouping result obtained by the grouping unit 316 through the output function of the input / output unit 317.
[0044] The storage unit 319 stores various data including various programs for operating the analysis system 3 and data generated by each unit (for example, segment length). The various programs also include programs such as waveform classification processing executed using a trained model. Furthermore, various data (including setting values) stored in the storage unit 319 can be dynamically reconfigured by the control unit 318. The storage unit 319 is configured using a ROM (Read Only Memory) preinstalled with various programs, a RAM (Random Access Memory) that stores calculation parameters and data for each process, an HDD (Hard Disk Drive), an SSD (Solid State Drive), and the like.
[0045] Various programs can also be recorded on computer-readable recording media such as HDDs, flash memory, CD-ROMs, DVD-ROMs, Blu-ray (registered trademark), and widely distributed, and the method therefor is not limited. Furthermore, the communication unit 311 can also acquire various programs via a communication network.
[0046] The analysis system 3 having the above functional configuration is a computer configured using one or more pieces of hardware such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), and an FPGA (Field Programmable Gate Array). The analysis system 3 may be configured separately from the electrocardiographic signal measuring device 2, or may be configured integrally with the electrocardiographic signal measuring device 2.
[0047] The receiving terminal 4 includes a communication unit 41, an output unit 42, and a control unit 43.
[0048] The communication unit 41 receives information output from the analysis system 3 via a communication network. For example, the communication unit 41 receives processing results and wearer information from the analysis system 3.
[0049] The output unit 42 is configured by an output device such as a display device or a printing device, and outputs various types of information under the control of the control unit 43. The output unit 42 has an output function provided by, for example, a display formed of liquid crystal, organic EL, or the like, or a speaker.
[0050] The control unit 43 integrally controls the operation of the receiving terminal 4. Further, the control unit 43 outputs the processing results by the output function of the output unit 42.
[0051] The receiving terminal 4 having the above functional configuration is a computer configured using one or more pieces of hardware such as a CPU, a GPU, an ASIC, and an FPGA. Note that the receiving terminal 4 may have a user interface.
[0052] The learning device 5 creates training data and generates a trained model using the created training data. In this embodiment, the learning performed by the learning device 5 can employ known machine learning methods such as Bayesian optimization, genetic algorithms, linear regression, nonlinear regression, logistic regression, k-nearest neighbors, support vector machines, decision trees, random forests, gradient boosting trees, k-means algorithms, principal component analysis, and deep learning. The trained model is, for example, a neural network consisting of an input layer, hidden layers, and output layers, with each layer having one or more nodes. Information such as network parameters in the trained model is stored in memory. Network parameters include information about the weights and biases between layers of the neural network. The learning device 5 generates the trained model. The learning device 5 is a computer composed of one or more hardware components such as a CPU, GPU, ASIC, FPGA, and memory.
[0053] Next, the waveform type classification process based on biological information obtained from the subject 1 will be explained. Figure 2 is a flowchart showing the waveform type classification process by the feature acquisition unit according to Embodiment 1. In the classification process, first, the analysis system 3 receives input of the electrocardiogram signal 10 obtained from the subject 1 (step S101). In this Embodiment 1, the input / output unit 317 is provided with an input port for a medium on which the electrocardiogram signal 10 is written, and after the acquisition of the electrocardiogram signal 10 is complete, the electrocardiogram signal 10 stored on the medium is read as an example, but the communication unit 311 may also receive input of the electrocardiogram signal 10. The control unit 318 processes the waveform data 11 of the acquired electrocardiogram signal 10 as needed to generate processing data.
[0054] After acquiring the electrocardiogram signal 10, the waveform detection unit 312 detects the electrocardiogram waveform from the input electrocardiogram signal (step S102). As described above, the electrocardiogram signal 10 may include waveform components called P wave, Q wave, R wave, S wave, and T wave. The waveform detection unit 312 can detect the electrocardiogram waveform included in the electrocardiogram signal by detecting these waveform components (especially the R wave).
[0055] After waveform detection, the feature acquisition unit 313 acquires multiple feature quantities for each detected electrocardiogram waveform (step 103).
[0056] Figure 3 is a flowchart showing the feature acquisition process by the feature acquisition unit according to Embodiment 1. First, the feature acquisition unit 313 applies a continuous wavelet transform as the time-frequency transformation of the electrocardiogram shape to acquire time-frequency representation data (step S201).
[0057] Figure 4 illustrates an example of time-frequency representation data (scalogram) obtained by applying a continuous wavelet transform to an electrocardiogram waveform. Figure 4(a) shows the waveform data before transformation. Figure 4(b) shows the time-frequency representation data, which is three-dimensional data with time (horizontal axis), frequency (vertical axis), and spectral intensity (shading). This three-dimensional data can be used for different waveform types because it expands the geometric characteristics of the electrocardiogram waveform in multiple dimensions.
[0058] Returning to Figure 3, the feature acquisition unit 313 acquires vectorized features by inputting the acquired time-frequency representation data into the trained model (step S202). In this embodiment 1, the trained model is a CNN (Convolutional Neural Network) model generated by the learning device 5. It extracts features from the input time-frequency representation data using a convolutional kernel in the convolutional layer and aggregates features by compressing information for each local space in the pooling layer. This makes it possible to utilize the time-frequency features of the time-frequency representation data as vectorized features.
[0059] In this case, it is assumed that the CNN model is trained using a general-purpose image dataset having the same structure as the input time-frequency representation data, but it is not limited to this, and models trained using time-frequency representation data or models obtained by performing transfer learning for time-frequency representation data are also assumed. The feature acquisition unit 313 reads the trained model from the learning device 5, for example, and acquires the features. If the generated trained model is stored in the storage unit 319, the feature acquisition unit 313 may read the trained model from the storage unit 319 and acquire the features.
[0060] Returning to Figure 2, the dimensionality reduction unit 314 constructs a dimensionality reduction model (step S104). The dimensionality reduction unit 314 combines the feature quantities acquired from each electrocardiogram waveform into one input and performs model construction of the UMAP, which is a dimensionality reduction model.
[0061] Here, a dimensionality reduction model is one that learns from multiple high-dimensional data inputs and enables efficient analysis and visualization by performing featureization (dimensionality reduction) in lower dimensions while preserving the structure and information of the data. UMAP is one such model. Compared to other dimensionality reduction models (for example, Principal Component Analysis (PCA)), UMAP excels at preserving the essential structure of high-dimensional data (local similarities between data and global relationships such as the distribution and patterns of the entire data) by constructing a neighborhood graph in high-dimensional space and projecting it onto a lower-dimensional space after optimization. This makes high-performance featureization possible even for data with complex nonlinear structures.
[0062] After model construction, the dimensionality reduction unit 314 performs featureization using the constructed dimensionality reduction model (step S105). The dimensionality reduction unit 314 features the features acquired from each electrocardiogram waveform in three dimensions using the constructed dimensionality reduction model and outputs the featureization results.
[0063] Figure 5 illustrates an example of a featureization result obtained by three-dimensionally characterizing the acquired features using UMAP. According to the featureization result shown in Figure 5, it can be considered that points in three dimensions corresponding to electrocardiogram waveforms that have similar features, i.e., points indicating similar waveform types, are visualized to be located at close coordinates in a three-dimensional space consisting of feature 1, feature 2, and feature 3.
[0064] Returning to Figure 2, the clustering unit 315 performs clustering (step S106). The clustering unit 315 performs nonlinear clustering on the output featureization results using the clustering model DBSCAN and outputs the clustering results. A clustering model, as referred to here, learns from multiple input data and classifies them into one or more clusters, thereby enabling data structuring, and DBSCAN is one such model. DBSCAN is a nonlinear clustering model that performs clustering based on data density, and compared to other linear clustering models, it enables visual clustering even for data sets with nonlinear and complex shapes. In nonlinear clustering using DBSCAN, the number of clusters is automatically determined by the algorithm. This clustering model is generated, for example, in the learning device 5.
[0065] Figure 6 is a diagram illustrating an example of clustering results obtained by clustering the featureization results using DBSCAN in the clustering unit 315. In Figure 6, electrocardiogram waveforms that are located at close coordinates in three-dimensional space, i.e., those that show similar waveform types, are clustered together and divided into a total of five clusters (clusters G1 to G5).
[0066] Figure 7 shows representative electrocardiogram waveforms in clusters clustered by the clustering unit according to Embodiment 1. In Figure 7, electrocardiogram waveforms with different characteristics, i.e., those exhibiting different waveform types (see, for example, the hatched areas), belong to different clusters. Thus, it can be seen that waveform types can be classified by utilizing the dimensionality reduction model and the clustering model.
[0067] Returning to Figure 2, the grouping unit 316 assigns a group number to each electrocardiogram waveform based on the clustering results and outputs the grouping results (step S107). By referring to the grouping results, the user can understand which group the electrocardiogram waveform belongs to.
[0068] Figure 8 is a diagram illustrating an example of the grouping results output by the grouping unit 316 according to Embodiment 1. As shown in Figure 8, a list is generated in which the time of occurrence of the electrocardiogram waveform (Index) and the group number (Group) are associated with the number (No) assigned to each electrocardiogram waveform. The grouping results are expected to be in a list format, such as shown in Figure 8, but are not limited to this. Furthermore, the information indicated by the Index may be in units of year, month, day, hour, minute, second, or it may be the number of samples taken since data acquisition began, and is not limited to this.
[0069] In this embodiment 1 described above, the features obtained from the electrocardiogram waveform are characterized using a dimensionality reduction model, and based on the characteristic results, similar waveforms are grouped together, thereby enabling high-precision classification of waveform types, such as arrhythmias. According to this embodiment 1, waveforms indicated by biological information can be classified with high precision.
[0070] (Embodiment 2) Next, Embodiment 2 of the present invention will be described. The biological information analysis system according to Embodiment 2 includes an analysis system 3A in place of the analysis system 3 according to Embodiment 1. The same reference numerals are used for the same components as described above.
[0071] Figure 9 shows an example of the configuration of the bio-information analysis system according to this second embodiment. The bio-information analysis system according to this second embodiment comprises an electrocardiogram signal measuring device 2, an analysis system 3A, a receiving terminal 4, and a learning device 5.
[0072] The analysis system 3A comprises a communication unit 311, a waveform detection unit 312, a feature acquisition unit 313, a dimensionality compression unit 314, a clustering unit 315, a grouping unit 316, a labeling unit 320, an input / output unit 317, a control unit 318, and a storage unit 319.
[0073] The labeling unit 320 performs waveform type labeling for each group of grouping results output from the grouping unit 316 and outputs the labeling results. According to the analysis system 3A, the labeling unit 320 enables labeling of each group based on the characteristic results of each labeled data.
[0074] Possible but not limited methods for labeling include selecting representative waveform data for each group, determining the waveform type based on features such as the generation interval and waveform shape of the selected representative waveform data, determining a label to indicate the waveform type, and applying the determined label to all waveform data in the group to which the representative waveform data belongs; or calculating statistical quantities related to features such as the generation interval and waveform shape of the waveform data for each group, determining a label based on these, and applying the determined label to all waveform data in that group.
[0075] In this second embodiment, labeled electrocardiogram waveform data is prepared in advance, and the process of referring to this labeled data and assigning labels to the waveform data based on the featureization results will be described as an example. In this case, it is assumed that the labeling unit 320 acquires the labeled data by having a labeled data input unit that accepts the input of labeled electrocardiogram waveform data, via the input / output unit 317, or by reading from the storage unit 319. Here, an example in which the labeled data input unit accepts the input of labeled electrocardiogram waveform data will be described.
[0076] Next, the waveform type classification process according to Embodiment 2 will be described. Figure 10 is a flowchart showing the waveform type classification process by the feature acquisition unit according to Embodiment 2. In the classification process according to Embodiment 2, the features obtained from the electrocardiogram waveform to be classified are characterized using a dimensionality reduction model, similar to steps S101 to S107 of Embodiment 1 described above (see Figure 2), and electrocardiogram waveforms showing similar waveform types are grouped together (steps S301 to S307).
[0077] After grouping, the labeled data input unit accepts the input of labeled electrocardiogram waveforms as data with labeled electrocardiogram waveforms (step S308). A labeled electrocardiogram waveform is an electrocardiogram waveform detected by the waveform detection unit 312 to which a label indicating the waveform type has been added.
[0078] Examples of labels indicating the waveform type of an electrocardiogram include, but are not limited to, N (normal), V (ventricular arrhythmia), S (atrial arrhythmia), and A (alteral conduction). Labeling can be done by, but are not limited to, visually determined labels assigned to the electrocardiogram waveform by a medical technologist, or by automatic determination based on features such as the interval between occurrences and waveform shape of the relevant electrocardiogram waveform. Furthermore, labeled electrocardiogram waveforms can be obtained from patients, from other samples, or generated by a device, but are not limited to these methods. Additionally, labeled data can be augmented using arbitrary data expansion or generation techniques and used as separate labeled data; the labels on the augmented data may be changed as appropriate.
[0079] In this second embodiment, it is assumed that the electrocardiogram waveform obtained from the patient is labeled N (normal) or V (ventricular arrhythmia) by a medical technologist based on visual inspection, and 10 data points for each label are prepared as labeled electrocardiogram waveforms and input into the labeled data input unit.
[0080] The feature acquisition unit 313 acquires multiple feature quantities for each waveform of the input labeled electrocardiogram waveform in the same manner as steps S201 to S202 shown in Figure 3 (step S309).
[0081] Then, the dimensionality reduction unit 314 performs only featureization on the features acquired from each labeled electrocardiogram waveform using the dimensionality reduction model constructed in step S304 (step S310).
[0082] Next, the labeling unit 320 performs labeling of each group based on the characteristic result of each labeled electrocardiogram waveform (step S311). After labeling, the grouping unit 316 updates and outputs the grouping results.
[0083] Figure 11 illustrates an example of the featureization results for each labeled electrocardiogram waveform used to label each group. In Figure 11, the clustering results for the electrocardiogram waveforms to be classified obtained in step S305 and the featureization results for the labeled electrocardiogram waveforms obtained in step S309 are superimposed and shown in three-dimensional space.
[0084] Based on the featureization results, electrocardiogram waveforms with similar features, i.e., similar waveform types, can be visualized in the three-dimensional space of Figure 11 so that they are located at close coordinates. For example, based on the coordinate information of labeled electrocardiogram shapes, the labels for each group can be determined.
[0085] Figures 12A to 12C are enlarged views of parts of the featureization results shown in Figure 11. Figure 12A is an enlarged view of Figure 11, focusing on the area around the cluster of group Gr1. Figure 12B is an enlarged view of the area around the cluster of group Gr2. Figure 12C is an enlarged view of the area around the cluster of group Gr3.
[0086] As shown in Figure 12A, electrocardiogram waveforms labeled N (normal) are located near the cluster of group Gr1. Also, as shown in Figure 12B, electrocardiogram waveforms labeled V (ventricular arrhythmia) are located near the cluster of group Gr2. Therefore, electrocardiogram waveforms belonging to group 1 can be labeled N (normal), and electrocardiogram waveforms belonging to group 2 can be labeled V (ventricular arrhythmia).
[0087] On the other hand, as shown in Figure 12C, near the cluster of group Gr3, electrocardiogram waveforms labeled N (normal) and electrocardiogram waveforms labeled V (ventricular arrhythmia) are mixed together. In this case, it is conceivable to adopt the label of the labeled electrocardiogram waveform closest to the cluster of group Gr3, or to adopt the label of the labeled electrocardiogram waveform with the highest number (density) in any given region.
[0088] The label for each group can be determined by the above procedure. In this embodiment 2, the labeling of each group was based on the coordinate information of the labeled electrocardiogram waveform, but it is not limited to this. In this embodiment 2, labels N (normal) and V (ventricular arrhythmia) were prepared as labeled electrocardiogram waveforms, but any form of label is acceptable as long as there is one or more labels. Furthermore, when performing labeling by the labeling unit, a label other than the labeled electrocardiogram waveform may be assumed, such as Unknown, which means that it does not fit any label.
[0089] According to Embodiment 2 described above, similar to Embodiment 1, the features obtained from the electrocardiogram waveform are characterized using a dimensionality reduction model, and based on the characteristic results, similar waveforms are grouped, thereby enabling high-precision classification of waveform types, such as arrhythmias. According to Embodiment 2, waveforms indicated by biological information can be classified with high precision.
[0090] Furthermore, according to this second embodiment, since each group is labeled based on the featureization results of each labeled data, a more detailed classification of waveform types can be performed.
[0091] (Embodiment 3) Next, Embodiment 3 of the present invention will be described. The biological information analysis system according to Embodiment 3 includes an analysis system 3B instead of the analysis system 3 according to Embodiment 1. The same reference numerals are used for the same components as described above.
[0092] Figure 13 shows an example of the configuration of a biological information analysis system according to this third embodiment. The biological information analysis system according to this third embodiment comprises an electrocardiogram signal measuring device 2, an analysis system 3B, a receiving terminal 4, and a learning device 5.
[0093] The analysis system 3B comprises a communication unit 311, a waveform detection unit 312, a feature acquisition unit 313, a dimensionality compression unit 314A, a clustering unit 315, a grouping unit 316, an input / output unit 317, a control unit 318, and a storage unit 319.
[0094] The dimensionality compression unit 314A includes a first dimensionality compression unit 314a and a second dimensionality compression unit 314b. The first dimensionality compression unit 314a characterizes a portion of the multiple features acquired by the feature acquisition unit 313 to a specified number of dimensions using a dimensionality compression model. The first dimensionality compression unit 314a outputs the characterized result as a first characterization result. The second dimensionality compression unit 314b characterizes any of the features acquired by the feature acquisition unit 313 to a specified number of dimensions using a dimensionality compression model, including at least the first characterization result. The second dimensionality compression unit 314b outputs the characterized result as a second characterization result.
[0095] In this embodiment 3, the first dimensionality compression unit 314a and the second dimensionality compression unit 314b are assumed to characterize in three dimensions using UMAP, which is one of the nonlinear dimensionality compression models, as the dimensionality compression model, but are not limited to this.
[0096] Next, the waveform type classification process according to Embodiment 3 will be described. Figure 14 is a flowchart showing the waveform type classification process by the feature acquisition unit according to Embodiment 3. In the classification process according to Embodiment 3, feature quantities are acquired from the electrocardiogram waveform to be classified in the same manner as steps S101 to S103 of Embodiment 1 described above (see Figure 2) (steps S401 to S403).
[0097] After acquiring the features, the first dimensionality compression unit 314a takes a portion of the acquired features as input and constructs a UMAP model as the first dimensionality compression model (step S404). In this embodiment 3, it is assumed, but is not limited to, that the features obtained in the same manner as in steps S201 to S202 shown in Figure 4 are used as input.
[0098] Then, the first dimensionality compression unit 314a characterizes the selected feature quantities in three dimensions using the first dimensionality compression model (step S405). The first dimensionality compression unit 314a outputs the characterization results obtained as the first characterization results to the second dimensionality compression unit 314b. In this embodiment 3, it is assumed that the feature quantities to be characterized by the first dimensionality compression unit 314a are feature vectors obtained in the same manner as steps S201 to S202 shown in Figure 4, but it is not limited to this, and for example, shape data such as waveform width, wave height, generation time, and generation time interval for each waveform component (P wave, Q wave, R wave, S wave, and T wave) of the waveform data are selected.
[0099] After the first feature acquisition result, the second dimensionality compression unit 314b takes some of the acquired features as input and constructs a UMAP model as the second dimensionality compression model (step S406).
[0100] In this embodiment 3, in addition to the first featureization result, which is a three-dimensional feature, it is assumed that the input will be a feature obtained by normalizing the ratio of the interval between the target electrocardiogram waveform and the preceding and succeeding electrocardiogram waveforms, but it is not limited to this. Furthermore, in this embodiment 3, although the first featureization result is a three-dimensional feature, some of the feature may be deleted to obtain a one-dimensional or two-dimensional feature before inputting it to the second dimension compression unit 314b.
[0101] Then, the second dimensionality compression unit 314b characterizes the selected features in three dimensions using the second dimensionality compression model (step S407). The second dimensionality compression unit 314b outputs the characterization results obtained through the characterization to the clustering unit 315 as the second characterization result.
[0102] Next, in the same manner as steps S106 and S107 of Embodiment 1 described above (see Figure 2), the clustering unit 315 performs nonlinear clustering on the second featureization result using the clustering model DBSCAN, and then the grouping unit 316 outputs the clustering result and the grouping result (steps S408 and S409).
[0103] According to Embodiment 3 described above, similar to Embodiment 1, the features obtained from the electrocardiogram waveform are characterized using a dimensionality reduction model, and based on the characteristic results, similar waveforms are grouped, thereby enabling high-precision classification of waveform types, such as arrhythmias. According to Embodiment 3, waveforms indicated by biological information can be classified with high precision.
[0104] Furthermore, in this embodiment 3, multiple feature quantities are characterized by a first dimensionality compression unit 314a, and then the resulting features are combined with other feature quantities and further characterized by a second dimensionality compression unit 314b. According to this embodiment 3, it is possible to perform multidimensional analysis using more feature quantities while increasing the interpretability of the feature quantities, thereby enabling highly accurate classification of waveform types.
[0105] (Embodiment 4) Next, Embodiment 4 of the present invention will be described. The biological information analysis system according to Embodiment 4 includes an analysis system 3C in place of the analysis system 3 according to Embodiment 1. The same reference numerals are used for the same components as described above.
[0106] Figure 15 shows an example of the configuration of the biological information analysis system according to this fourth embodiment. The biological information analysis system according to this second embodiment comprises an electrocardiogram signal measuring device 2, an analysis system 3C, a receiving terminal 4, and a learning device 5.
[0107] The analysis system 3C comprises a communication unit 311, a waveform detection unit 312, a screening unit 321, a feature acquisition unit 313, a dimensionality compression unit 314, a clustering unit 315, a grouping unit 316, an input / output unit 317, a control unit 318, and a storage unit 319.
[0108] The screening unit 321 extracts waveform data to be classified from the waveform data detected by the waveform detection unit according to the set rules.
[0109] Next, the waveform type classification process according to Embodiment 4 will be described. Figure 16 is a flowchart showing the waveform type classification process by the feature acquisition unit according to Embodiment 4. In the classification process according to Embodiment 4, the electrocardiogram waveform is detected in the same manner as in steps S101 and S102 of Embodiment 1 described above (see Figure 2) (steps S501 and S502).
[0110] Next, the screening unit 321 extracts electrocardiogram waveforms to be classified from the detected electrocardiogram patterns according to the set rules (step S503). Specific examples of extracting electrocardiogram waveforms to be classified based on the rules include extracting electrocardiogram waveforms with a amplitude above a threshold based on the amplitude of the electrocardiogram pattern, extracting high-quality electrocardiogram waveforms based on the amount of noise superimposed on the electrocardiogram waveform, and excluding clearly normal waveforms and extracting electrocardiogram waveforms that raise concerns about arrhythmias based on occurrence intervals, waveform shape, etc. According to these rules, by excluding elements unnecessary for the purpose of classifying waveform types, such as features caused by the measurement conditions (fluctuations in amplitude and superimposed noise) and features of normal waveforms of low importance for classification, it is possible to focus on and characterize the waveforms in a more essential way.
[0111] Next, in the analysis system 3C, in the same manner as in steps S103 to S107 of Embodiment 1 described above (see Figure 3), the feature quantities obtained from the electrocardiogram waveforms extracted in the screening unit are characterized using a dimensionality reduction model, electrocardiogram waveforms showing similar waveform types are grouped together, and the grouping results are output (steps S504 to S508).
[0112] According to Embodiment 4 described above, similar to Embodiment 1, the features obtained from the electrocardiogram waveform are characterized using a dimensionality reduction model, and based on the characteristic results, similar waveforms are grouped, thereby enabling high-precision classification of waveform types, such as arrhythmias. According to Embodiment 4, waveforms indicating biological information can be classified with high precision.
[0113] Furthermore, according to this embodiment 4, by extracting waveform data to be classified according to the set rules, unnecessary elements for the purpose of classifying waveform types, such as features caused by the measurement conditions (fluctuations in wave height and superposition of noise) and features of normal waveforms that are of low importance for classification, are excluded. This allows for characterization that focuses on the more essential structure.
[0114] (Embodiment 5) Next, Embodiment 5 of the present invention will be described. The biological information analysis system according to Embodiment 5 includes an analysis system 3D in place of the analysis system 3 according to Embodiment 1. The same reference numerals are used for the same components as described above.
[0115] Figure 17 shows an example of the configuration of the biological information analysis system according to this embodiment 5. The biological information analysis system according to this embodiment 2 comprises an electrocardiogram signal measuring device 2, an analysis system 3D, a receiving terminal 4, and a learning device 5.
[0116] The analysis system 3D comprises a communication unit 311, a waveform detection unit 312, a feature acquisition unit 313, an extraction unit 322, a dimensionality compression unit 314, a clustering unit 315, a grouping unit 316, an input / output unit 317, a control unit 318, and a storage unit 319.
[0117] The extraction unit 322 extracts waveform data belonging to a specified group and feature quantities related to that waveform data, based on the grouping results output by the grouping unit 316 or as an additional analysis.
[0118] In the 3D analysis system, additional analysis enables featureization that utilizes features that were not considered or could not be considered in the overall data by focusing on specific groups and performing re-featuration. In this embodiment 5, the dimensionality reduction of the additional analysis is assumed to be performed using UMAP, a nonlinear dimensionality reduction model, as the dimensionality reduction model, to feature the data in three dimensions, and the clustering of the additional analysis is assumed to be performed using DBSCAN, a nonlinear clustering model, as the clustering model, but the system is not limited to these.
[0119] Next, the waveform type classification process according to Embodiment 5 will be described. Figure 18 is a flowchart showing the waveform type classification process by the feature acquisition unit according to Embodiment 5. In the classification process according to Embodiment 5, the features obtained from the electrocardiogram waveform to be classified are characterized using a dimensionality reduction model, similar to steps S101 to S107 of Embodiment 1 described above (see Figure 2), and electrocardiogram waveforms showing similar waveform types are grouped together (steps S601 to S607).
[0120] Then, the extraction unit 322, based on the grouping results output from the grouping unit 316, selects a group as the group to be further analyzed and extracts the electrocardiogram waveforms and their associated feature quantities (step S608). This process may be applied not only to the grouping results output from the grouping unit 316, but also to the grouping results output as additional analysis, which will be described later.
[0121] The additional analysis target group here may be a single group or multiple groups combined. Furthermore, the additional analysis target group may include all groups or only those groups that satisfy the set rules. Specific examples of specifying the additional analysis target group according to the rules include specifying a group whose number of elements in the electrocardiogram shape is greater than or equal to a threshold, specifying a group whose difference from other groups in the three-dimensional space of the featureization results is greater than or equal to a threshold, or specifying a group whose group variability in the three-dimensional space of the featureization results is greater than or equal to a threshold. In this embodiment 5, it is assumed that only the group with group number 1 in the grouping results of step S607 will be targeted as the additional analysis target group.
[0122] Once the feature quantities of the additional analysis target group are extracted, the control unit 318 performs additional analysis on the feature quantities of the additional analysis group. First, the dimensionality reduction unit 314 using additional analysis takes the extracted electrocardiogram waveform feature quantities as input and constructs a UMAP model, which is a dimensionality reduction model (step S609).
[0123] Then, the dimensionality reduction unit 314 characterizes the selected features in three dimensions using the constructed dimensionality reduction model (step S610). The dimensionality reduction unit 314 outputs the characterization results obtained through the characterization process to the clustering unit 315 as the characterization results.
[0124] Next, the clustering unit 315, which performs additional analysis, performs nonlinear clustering on the output featureization results using the clustering model DBSCAN (step S611). The clustering unit 315 outputs the clustering results to the grouping unit 316.
[0125] Next, the grouping unit 316, based on the additional analysis, updates the group number for each electrocardiogram waveform based on the clustering results (step S612). The grouping unit 316 outputs the new grouping results.
[0126] Figure 19 is a diagram illustrating an example of clustering results obtained by the clustering unit 315 in the clustering of the additional analysis in Embodiment 5. In Figure 19, the additional analysis target groups are assumed to be group Gr11 with group number 1, group Gr12 with group number 2, and group Gr13 with group number 3. For example, in Figure 6, group Gr11 is clustered together as one cluster G1 in three-dimensional space, but in Figure 19, it can be seen that it is separated into three clusters in three-dimensional space. This is because, by focusing the characterization on the additional analysis target groups in step S610, it became possible to utilize features that were not considered or could not be considered in step S605.
[0127] Figure 20 shows representative electrocardiogram waveforms for each of the three clusters clustered by the clustering unit 315 of Embodiment 5. Figure 20 shows representative electrocardiogram waveforms for each of the three clusters clustered in Figure 19. In Figure 20, electrocardiogram waveforms with different characteristics, i.e., different waveform types, belong to different clusters, and it can be seen that it is possible to further subdivide the classification of waveform types by additional analysis.
[0128] Next, the control unit 318 determines whether there are any unupdated additional analysis target groups (step S613). For example, the control unit 318 extracts groups from the additional analysis target groups that have not undergone processing such as feature extraction, and if the number of such groups is one or more, it determines that there are unupdated groups. If the control unit 318 determines that there are unupdated additional analysis target groups, it repeats the processing in steps S608 to S612 for those unupdated additional analysis target groups. Note that groups whose group numbers have been updated in step S612 may be designated as additional analysis target groups again, and in this case, an upper limit on the number of repetitions may be set for each group / waveform data. Furthermore, the feature quantities extracted by the extraction unit 322, the rules for specifying additional analysis target groups in the extraction unit 322, the specifications of the dimensionality compression model in the dimensionality compression of the additional analysis by the dimensionality compression unit 314, the specifications of the clustering model in the clustering unit 315 by the additional analysis, and the specifications of the grouping results in the grouping unit 316 by the additional analysis can be changed as appropriate.
[0129] According to Embodiment 5 described above, similar to Embodiment 1, the features obtained from the electrocardiogram waveform are characterized using a dimensionality reduction model, and based on the characteristic results, similar waveforms are grouped, thereby enabling high-precision classification of waveform types, such as arrhythmias. According to Embodiment 5, waveforms indicated by biological information can be classified with high precision.
[0130] Furthermore, according to this embodiment 5, by performing additional analysis by narrowing the focus to groups and re-characterizing the data, features that were not noticed or could not be noticed in the entire data are utilized for characterization, thus enabling the acquisition of even more detailed classification results for waveform types.
[0131] (Other Embodiments) While embodiments for carrying out the present invention have been described so far, the present invention should not be limited to the embodiments described above. For example, although the analysis system has been described as being provided separately from the electrocardiogram signal measuring device, the analysis system may be configured integrally with the electrocardiogram signal measuring device.
[0132] Furthermore, each process, such as the waveform detection process by the waveform detection unit 312, the feature acquisition process by the feature acquisition unit 313, the dimensionality compression process by the dimensionality compression units 314 and 314A, the dimensionality compression process by the first dimensionality compression unit 314a and the second dimensionality compression unit 314b, the clustering process by the clustering unit 315, the grouping process by the grouping unit 316, the labeling process by the labeling unit 320, the screening process by the screening unit 321, and the extraction process by the extraction unit 322, is preferably, but not limited to, being performed mechanically based on the command sequence of the control unit 318 and the setting values stored in the storage unit 319, thereby eliminating human judgment.
[0133] 1 Subject 2 Electrocardiogram signal measuring device 3, 3A-3D Analysis system 4 Receiving terminal 5 Learning device 10 Electrocardiogram signal 11 Waveform data 41, 311 Communication unit 42 Output unit 43, 318 Control unit 100, 110, 120 Electrocardiogram waveform 312 Waveform detection unit 313 Feature acquisition unit 314, 314A Dimensional compression unit 314a First dimensional compression unit 314b Second dimensional compression unit 315 Clustering unit 316 Grouping unit 317 Input / Output unit 319 Storage unit 320 Labeling unit 321 Screening unit 322 Extraction unit
Claims
1. A biological information analysis system that classifies data using target biological information, comprising: an input receiving unit that receives input of the biological information; a waveform detection unit that detects waveform data from the biological information; a feature acquisition unit that acquires a plurality of feature quantities relating to the waveform data; a dimensionality reduction unit that characterizes the feature quantities using a dimensionality reduction model; a clustering unit that clusters the featureization results using a clustering model; and a grouping unit that groups the waveform data based on the clustering results.
2. The biological information analysis system according to claim 1, wherein the biological information is information indicating the activity of the target heart, and the biological information is used to classify arrhythmias.
3. The biological information analysis system according to claim 1, further comprising a screening unit for extracting waveform data to be classified from the waveform data.
4. The biological information analysis system according to claim 1, wherein the dimensionality reduction model is a nonlinear dimensionality reduction model.
5. The bio-information analysis system according to claim 1, wherein the clustering model is a nonlinear clustering model.
6. The biological information analysis system according to claim 1, wherein the feature quantities include at least one of the following: time-frequency representation data obtained by applying a time-frequency transformation to the waveform data; shape data obtained from the width of the waveform components in the waveform data; and periodic data relating to the interval between preceding and succeeding waveform data in the waveform data.
7. The biological information analysis system according to claim 1, further comprising: an extraction unit that extracts waveform data belonging to a specified group based on the grouping results; and a control unit that performs featureization by the dimensionality compression unit, clustering by the clustering unit, and grouping by the grouping unit on the waveform data belonging to the specified group.
8. The biological information analysis system according to claim 1, characterized in that the dimensionality reduction model characterizes the feature quantities in two dimensions, and the characterization result is two-dimensional numerical data.
9. The biological information analysis system according to claim 1, characterized in that the dimensionality reduction model characterizes the feature quantities in three dimensions, and the characterization result is three-dimensional numerical data.
10. The biological information analysis system according to claim 1, further comprising a labeling unit that labels each group with a waveform type based on the grouping results.
11. The biological information analysis system according to claim 10, wherein the labeling unit selects a representative waveform data from the waveform data belonging to each group, determines the waveform type of the representative waveform data, and labels all the waveform data of the group to which the representative waveform data belongs with the waveform type.
12. The biological information analysis system according to claim 10, wherein the labeling unit determines a waveform type based on a statistical amount calculated for each group and labels all waveform data belonging to each group with the waveform type.
13. The biological information analysis system according to claim 10, further comprising a labeled data input unit that accepts labeled waveform data, wherein the labeling unit characterizes the feature quantities obtained from the labeled waveform as labeled waveform data using the dimensionality reduction model to obtain labeled featureization results, and labels the waveform type for each group based on the labeled featureization results.
14. The dimensionality reduction unit comprises a first dimensionality reduction unit and a second dimensionality reduction unit, the first dimensionality reduction unit obtains a first featureization result by characterizing a part of the feature quantities using a dimensionality reduction model, the second dimensionality reduction unit obtains a second featureization result by characterizing any of the feature quantities using a dimensionality reduction model, including at least the first featureization result, and the clustering unit clusters the second featureization result using a clustering model, the biological information analysis system according to claim 1.
15. A bio-information analysis program that classifies data using target bio-information, wherein the program causes a computer to perform a series of processes, including: receiving input of the bio-information; detecting waveform data from the bio-information; obtaining a plurality of feature quantities related to the waveform data; characterizing the feature quantities using a dimensionality reduction model; clustering the featureization results using a clustering model; and grouping the waveform data based on the clustering results.