Method and apparatus for diagnosing diseases by using artificial intelligence model
An AI-driven method processes Raman signal data to diagnose diseases like pancreatic cancer, overcoming the lack of effective biomarkers in existing spectroscopy methods, ensuring accurate and repeatable results.
Patent Information
- Application Number
- PCT/KR2025/007021
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-05-22
- Filing Date
- 2025-05-23
- Publication Date
- 2025-12-04
AI Technical Summary
Existing surface-enhanced Raman scattering spectroscopy methods struggle with the lack of effective biomarkers for diagnosing intractable diseases like pancreatic cancer, making screening and early diagnosis difficult.
A method and device using an artificial intelligence model that processes Raman signal data from biological solutions, pre-processes the data, and inputs it into a diagnostic model to generate disease information vectors, enabling disease diagnosis without relying on specific biomarkers.
Enables accurate disease diagnosis with high repeatability and convenience for distribution and storage, providing diagnostic results for various diseases, including pancreatic cancer, using surface-enhanced Raman scattering spectroscopy.
Smart Images

Figure KR2025007021_04122025_PF_FP_ABST
Abstract
Description
Method and device for diagnosing disease using an artificial intelligence model
[0001] The present disclosure relates to a method and device for diagnosing a disease using an artificial intelligence model.
[0002] Surface-enhanced Raman scattering (SERS) spectroscopy is a spectroscopic technique designed to complement Raman scattering spectroscopy, which has weak signals and low reproducibility. It is a spectroscopic technique in which the Raman scattering intensity of molecules adsorbed on the surface of metal nanostructures such as gold and silver increases rapidly by 10 6 ~ 10 8 It is a spectroscopy method that utilizes the phenomenon of increasing by more than a factor of two.
[0003] Surface-enhanced Raman scattering (SERS) spectroscopy is a technique that can obtain a large amount of information in a single measurement, is ultrasensitive enough to directly measure a single molecule, and can directly measure information about the vibrational state or molecular structure of a molecule, so it is recognized as a powerful analytical method for chemical / biological / biochemical analysis.
[0004] Recently, surface-enhanced Raman scattering spectroscopy has been attracting attention as a method for precisely obtaining comprehensive information on biomolecules from bioanalytical samples. However, the detection method for biomolecules using surface-enhanced Raman scattering spectroscopy proposed in the past requires specific binding, and there was a problem that screening and early diagnosis were difficult because there was no biomarker with excellent diagnostic performance among blood indicators for intractable cancers and intractable diseases such as pancreatic cancer.
[0005] Accordingly, there is a need for a diagnostic method and diagnostic system using surface-enhanced Raman scattering spectroscopy that can be used even when there is no biomarker with excellent diagnostic performance among biomarkers, such as intractable cancers and incurable diseases such as pancreatic cancer.
[0006] Accordingly, the present invention aims to diagnose various diseases, not just pancreatic cancer, using surface-enhanced Raman scattering spectroscopy and an artificial intelligence model.
[0007] The present invention provides a method and device for diagnosing diseases using an artificial intelligence model. Furthermore, the present invention provides a computer-readable recording medium containing a program for executing the method on a computer. The technical challenges to be addressed are not limited to the technical challenges described above, and other technical challenges may also exist.
[0008] According to one aspect of the present disclosure, a method for diagnosing a disease using an artificial intelligence model may be provided, including: a step of obtaining a plurality of Raman signal data from a biological solution of a subject using a Raman signal intensifier; a step of performing a predetermined preprocessing on the plurality of Raman signal data to generate a plurality of Raman signal derived data corresponding to each of the plurality of Raman signal data; a step of inputting the plurality of Raman signal derived data into a diagnostic model as input data and obtaining a plurality of disease information vectors corresponding to each of the plurality of Raman signal derived data as output data; and a step of obtaining disease information of the subject based on the plurality of disease information vectors.
[0009] According to another aspect of the present disclosure, a device includes a memory storing at least one program; and at least one processor executing the at least one program; wherein the at least one processor obtains a plurality of Raman signal data from a biological solution of a subject using a Raman signal amplifier, performs a predetermined preprocessing on the plurality of Raman signal data to generate a plurality of Raman signal derived data corresponding to each of the plurality of Raman signal data, inputs the plurality of Raman signal derived data into a diagnostic model as input data, and obtains a plurality of disease information vectors corresponding to each of the plurality of Raman signal derived data as output data, and can obtain disease information of the subject based on the plurality of disease information vectors.
[0010] A computer-readable recording medium according to another aspect of the present disclosure includes a recording medium having recorded thereon a program for executing the above-described method on a computer.
[0011] According to one embodiment of the present invention, even if there is no biomarker with excellent diagnostic performance among biomarkers, a disease can be diagnosed using surface-enhanced Raman scattering spectroscopy.
[0012] In addition, surface-enhanced Raman scattering spectroscopy, which has high repeatability and is convenient for distribution and storage, can be used to provide accurate diagnostic results to patients.
[0013] The purposes of the present invention are not limited to the purposes mentioned above, and other purposes not mentioned will be clearly understood by those skilled in the art from the description below.
[0014] FIG. 1 is a diagram illustrating an example of a method for diagnosing a disease using an artificial intelligence model according to one embodiment.
[0015] FIG. 2 is a block diagram illustrating an example of a device for diagnosing a disease using an artificial intelligence model according to one embodiment.
[0016] FIG. 3 is a flowchart illustrating an example of a method for diagnosing a disease using an artificial intelligence model according to one embodiment.
[0017] FIGS. 4A to 4D are drawings for explaining an example of a structure for surface-enhanced Raman scattering spectroscopy according to one embodiment.
[0018] FIG. 5 is a drawing for explaining an example of a composition for surface-enhanced Raman scattering spectroscopy according to one embodiment.
[0019] FIG. 6 is a flowchart illustrating an example of a method for generating multiple Raman signal derivative data according to one embodiment.
[0020] FIG. 7 is a diagram illustrating an example of a method for obtaining multiple disease information vectors according to one embodiment.
[0021] FIG. 8 is a diagram illustrating an example of a method for obtaining disease information of a subject based on a plurality of disease information vectors according to one embodiment.
[0022] FIG. 9 is a diagram illustrating an example of a method for learning a diagnostic model for diagnosing a disease according to one embodiment.
[0023] FIG. 10 is a drawing for explaining an example of a method for obtaining disease information of a subject according to one embodiment.
[0024] A device for diagnosing a disease using an artificial intelligence model according to one aspect comprises: at least one memory; and at least one processor; wherein the at least one processor obtains a plurality of Raman signal data from a biological solution of a subject using a Raman signal amplifier, performs a predetermined preprocessing on the plurality of Raman signal data to generate a plurality of Raman signal derived data corresponding to each of the plurality of Raman signal data, inputs the plurality of Raman signal derived data into a diagnostic model as input data, obtains a plurality of disease information vectors corresponding to each of the plurality of Raman signal derived data as output data, and obtains disease information of the subject based on the plurality of disease information vectors.
[0025] The terms used in the examples are selected from widely used, current terms, as much as possible. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, the applicant may arbitrarily select terms, in which case their meanings will be described in detail in the relevant description. Therefore, the terms used in the specification should be defined based on their intended meaning and the overall content of the specification, rather than simply their names.
[0026] When a part of the specification is said to "include" a component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "unit" and "module" used throughout the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.
[0027] Additionally, terms including ordinal numbers, such as "first" or "second," used in the specification may be used to describe various components, but the components should not be limited by the terms. The terms may be used to distinguish one component from another.
[0028] The present disclosure will now be described in detail with reference to the attached drawings. Specifically, a method for diagnosing a disease using an artificial intelligence model according to one embodiment will be described in more detail with reference to FIGS. 1 through 10. However, the embodiments may be implemented in various different forms and are not limited to the examples described herein.
[0029] FIG. 1 is a diagram illustrating an example of a method for diagnosing a disease using an artificial intelligence model according to one embodiment.
[0030] Hereinafter, with reference to Fig. 1, an example of a method for diagnosing a disease using an artificial intelligence model is described.
[0031] Referring to Fig. 1, a patient with a disease (1) can be provided with disease information (100) obtained using an artificial intelligence model (10).
[0032] For example, a biomolecular material can be obtained from a patient with a disease (1), and a signal for diagnosing the disease of the patient (1) can be obtained using the obtained biomolecular material. In addition, the signal for diagnosing the disease of the patient (1) can be input as input data to an artificial intelligence model (10), and disease information (100) of the patient (1) can be obtained as output data.
[0033] For example, a signal for diagnosing a disease of a patient (1) may be a Raman signal. Here, the Raman signal may be a signal acquired based on Raman spectroscopy, and a method for acquiring a Raman signal using a biomolecular material of a patient (1) will be described below with reference to FIGS. 4a to 5.
[0034] Here, the subject from which disease information (100) can be obtained may be animals, not just humans. In addition, the disease information (100) may include, but is not limited to, information on degenerative brain diseases such as Alzheimer's disease and Parkinson's disease, cancers such as pancreatic cancer, thyroid cancer, prostate cancer, breast cancer, stomach cancer, and colon cancer, and infectious diseases such as hepatitis, coronavirus, and influenza.
[0035]
[0036] FIG. 2 is a block diagram illustrating an example of a device for diagnosing a disease using an artificial intelligence model according to one embodiment.
[0037] Referring to FIG. 2, a device (hereinafter referred to as "device") (200) for diagnosing a disease using an artificial intelligence model may include a communication unit (210), a processor (220), and a memory (230). Only components related to the embodiment are illustrated in the device (200) of FIG. 2. Therefore, it is apparent to those skilled in the art that other general components may be included in addition to the components illustrated in FIG. 2.
[0038] The communication unit (210) may include one or more components that enable wired / wireless communication with an external server or external device. For example, the communication unit (210) may include a short-range communication unit (not shown) and a mobile communication unit (not shown) for communication with an external server or external device.
[0039] For example, the device (200) can be connected to an external server via wired or wireless communication to transmit / receive data (e.g., clinical information of a patient, disease information of a patient, clinical information of a subject, disease information of a subject, Raman signal data, Raman signal derived data, etc.) between the two.
[0040] For example, the external server may be a computing device (e.g., a cloud) that includes separate memory and processor and has its own computing capabilities. Accordingly, at least some of the operations performed by the processor (220) described below may also be performed by the external server. In other words, at least some of the operations of the processor (220) described with reference to FIGS. 3 to 10 may be performed by the external server.
[0041] The memory (230) is hardware that stores various data processed within the device (200), and can store a program for processing and controlling the processor (220).
[0042] For example, the memory (230) may store various data such as Raman signal data, Raman signal derived data, disease information of a patient, clinical information of a patient, disease information of a subject, clinical information of a subject, and data generated according to the operation of the processor (220). In addition, the memory (230) may store an operating system (OS) and at least one program (e.g., a program required for the processor (220) to operate).
[0043] The memory (230) may include random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray or other optical disk storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.
[0044] The processor (220) controls the overall operation of the device (200). For example, the processor (220) can control the input unit (not shown), the display (not shown), the communication unit (210), the memory (230), etc., by executing programs stored in the memory (230).
[0045] The processor (220) may be implemented using at least one of application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controller units, microprocessors, and other electrical units for performing functions.
[0046] The processor (220) can control the operation of the device (200) by executing programs stored in the memory (230). For example, the processor (220) can perform at least a part of the method for diagnosing a disease using an artificial intelligence model, described with reference to FIGS. 3 to 10 . In addition, as described above, an external server can also perform at least a part of the method for diagnosing a disease using an artificial intelligence model, described with reference to FIGS. 3 to 10 .
[0047]
[0048] FIG. 3 is a flowchart illustrating an example of a method for diagnosing a disease using an artificial intelligence model according to one embodiment.
[0049] Hereinafter, with reference to FIG. 3, an example of a method for diagnosing a disease using an artificial intelligence model is described.
[0050] Referring to FIG. 3, a method for diagnosing a disease using an artificial intelligence model may include steps 310 to 340. However, the present invention is not limited thereto, and other general operations may be further included in the method for diagnosing a disease using an artificial intelligence model in addition to the operations illustrated in FIG. 3. Furthermore, as described above with reference to FIGS. 1 and 2, at least one of the operations in the flowchart illustrated in FIG. 3 may be processed by the processor (220).
[0051] First, in step 310, the processor (220) can acquire a plurality of Raman signal data from the biological solution of the subject using a Raman signal intensifier. Here, the Raman signal data may be data acquired by irradiating light on a material including the Raman signal intensifier.
[0052] For example, the processor (220) can obtain multiple Raman signal data using a single biological solution obtained from a subject using a Raman signal intensifier. That is, the processor (220) can obtain multiple Raman signal data from the same biological solution.
[0053] For example, a Raman signal enhancer may be a material that enables effective acquisition of a Raman signal and may include a metal nanostructure. More specifically, it may be a material that includes a plasmonic metal (or plasmonic nanoparticle), and details are described below with reference to FIGS. 4A to 4D.
[0054] In step 320, the processor (220) may perform a predetermined preprocessing on a plurality of Raman signal data to generate a plurality of Raman signal derivative data corresponding to each of the plurality of Raman signal data.
[0055] For example, when multiple Raman signal data are acquired using one biological solution obtained from a subject, the processor (220) can generate multiple Raman signal derivative data corresponding to each of the multiple Raman signal data.
[0056] Specifically, the processor (220) can extract some Raman spectrum data related to a specific disease from among the Raman spectrum data included in each of the plurality of Raman signal data based on the size of the Raman signal, normalize some of the Raman spectrum data, and image the normalized some of the Raman spectrum data through a predetermined preprocessing to generate a plurality of Raman signal derived data.
[0057] In step 330, the processor (220) can input a plurality of Raman signal derived data as input data into the diagnostic model and obtain a plurality of disease information vectors corresponding to each of the plurality of Raman signal derived data as output data.
[0058] For example, the processor (220) may input a plurality of Raman signal derived data into the diagnostic model as input data of the diagnostic model, calculate a probability value that the subject has at least one disease for each of the plurality of Raman signal derived data, and obtain a plurality of disease information vectors including the probability values as output data of the diagnostic model.
[0059] Meanwhile, the diagnostic model in the present disclosure may mean a single diagnostic model when multiple Raman signal data are acquired using a single biological solution obtained from a subject, or at least one of a first diagnostic model and a second diagnostic model when a single Raman signal data is acquired using a single biological solution. In addition, the diagnostic model may mean an artificial intelligence model trained to detect a disease of a subject using the patient's Raman signal data as learning data.
[0060] Additionally, in step 340, the processor (220) can obtain disease information of the subject based on a plurality of disease information vectors.
[0061] For example, the processor (220) may calculate an average value of the probability values that the subject has at least one disease included in each of a plurality of disease information vectors, for at least one disease, and obtain disease information of the subject based on the average value calculated for at least one disease.
[0062]
[0063] Again, referring to FIG. 3, in step 310, the processor (220) can acquire a plurality of Raman signal data from the biological solution of the subject using a Raman signal intensifier.
[0064] Below, an example of obtaining Raman signal data is described in detail.
[0065] FIGS. 4A to 4D are drawings for explaining an example of a structure for surface-enhanced Raman scattering spectroscopy according to one embodiment.
[0066] Referring to FIG. 4a, a surface-enhanced Raman scattering spectroscopy structure (400) (hereinafter referred to as “structure”) may include nanoparticles (P) and a silica shell layer (430).
[0067] In addition, as an optional embodiment, a biomolecular material (B) to be described later may be included within the silica shell layer (430), and a surface-enhanced Raman scattering spectral signal (or Raman signal) may be obtained for a significant number of biomolecular materials (B) non-specifically disposed within the structure (400) without excluding specific biomolecular materials (B) through the structure. In the following, the spectral signal or Raman signal used may be interpreted to have the same meaning.
[0068] Furthermore, by using a structure (400) like this and a composition for surface-enhanced Raman scattering spectroscopy in which these structures (400) are dispersed, by creating a database of surface-enhanced Raman scattering signals, a diagnostic kit having excellent diagnostic performance can be provided even for diseases without biomarkers with excellent spectral performance. In addition, the databased signals can be stored in a memory (230) or a separate storage device.
[0069] Below, the structure (400) is described.
[0070] In one embodiment, the nanoparticle (P) may include a plasmonic metal. The surface-enhanced Raman scattering spectroscopy signal can be obtained by increasing the Raman scattering signal through a plasmon phenomenon at a local surface when molecules or particles are present near the surface of the plasmonic metal.
[0071] As used herein, the term “plasmonic metal” may refer to a metal that is excited in a magnetic field, and specifically may refer to, but is not limited to, Au, Ag, Cu, Al, W, Pt, Ni, and Pd.
[0072] In one embodiment, the nanoparticle (P) may include a core (410) including a first plasmonic metal, and a magnetic shell layer (20) disposed to surround the core (410) and including a second plasmonic metal. In this way, by allowing the nanoparticle (P) to have a core-shell structure, it is possible to provide a nanoparticle (P) having both properties of the first plasmonic metal and the second plasmonic metal.
[0073] In one embodiment, the magnetic shell layer (420) may have a symmetrical structure based on the center (M) of the magnetic shell layer (420). In this way, by having the magnetic shell layer (420) surround the core (410) and have a symmetrical structure, the surface area between the nanoparticle (P) and the biomolecular material (B) can be increased, thereby maximizing the surface-enhanced Raman scattering spectroscopy signal.
[0074] In one embodiment, the magnetic shell layer (420) may have a regular polyhedral structure, and specifically may include at least one of a regular cube, a regular octahedron, a regular dodecahedron, and a regular icosahedron, but is not limited thereto.
[0075] In one embodiment, the core (410) includes a first plasmonic metal, and the first plasmonic metal is any metal that can provide a surface-enhanced Raman scattering spectroscopy signal through a plasmon phenomenon on the surface, and is interpreted to fall within the scope of the present invention, and examples thereof include, but are not limited to, Au, Ag, Cu, Al, W, Pt, Ni, and Pd.
[0076] In one embodiment, the core (410) may have an average particle diameter (D50) of 50 nm or less. As another example, 0 to 50, 5 to 50, 10 to 50, 15 to 50, 20 to 50, 25 to 50, 30 to 50, 35 to 50, 40 to 50, 45 to 50, 0 to 45, 5 to 45, 10 to 45, 15 to 45, 20 to 45, 25 to 45, 30 to 45, 35 to 45, 40 to 45, 0 to 40, 5 to 40, 10 to 40, 15 to 40, 20 to 40, 25 to 40, 30 to 40, 35 to 40, 0 to 35, 5 to 35, 10 to 35, 15 to It can be 35, 20 to 35, 25 to 35, 30 to 35, 0 to 30, 5 to 30, 10 to 30, 15 to 30, 20 to 30, 25 to 30, 0 to 25, 5 to 25, 10 to 25, 15 to 25, or 20 to 25.
[0077] When the structures have an average particle size as described above and are arranged in a composition (e.g., dispersion) to be described later, an excellent Raman scattering effect can be provided.
[0078] In one embodiment, the core (410) may have a diameter of 10 nm to 50 nm in at least one direction. By having such a structure, when the structures (400) are arranged in the composition described below, an excellent Raman scattering effect can be provided.
[0079] In one embodiment, the magnetic shell layer (420) includes a second plasmonic metal, and the width of the plasmon resonance peak on the surface of the second plasmonic metal can be narrower than the width of the plasmon resonance peak on the surface of the first plasmonic metal. By making the width of the plasmon resonance peak on the surface of the second plasmonic metal narrower, such as in this structure, when the structure (400) is dispersed in the composition described below, better sensitivity can be provided when acquiring a Raman scattering signal.
[0080] Here, as an optional embodiment, the width of the plasmon resonance peak on the surface of the first plasmonic metal and the width of the plasmon resonance peak on the surface of the second plasmonic metal can be measured in the same environment.
[0081] In one embodiment, the magnetic shell layer (420) includes a second plasmonic metal, and the second plasmonic metal includes, but is not limited to, Au, Ag, Cu, Al, W, Pt, Ni, and Pd.
[0082] In one embodiment, the core (410) may include gold (Au), and the magnetic shell layer (420) may include silver (Ag). Silver (Ag) has a light-gathering effect that is at least five times stronger than gold (Au), and a structure (400) including nanoparticles (P) having such a structure may provide an excellent surface-enhanced Raman scattering spectroscopy signal.
[0083] In one embodiment, the magnetic shell layer (420) may have a thickness of 10 nm to 70 nm in at least one direction. As another example, the thickness of the magnetic shell layer (420) may be 30 to 70 nm, 40 to 70 nm, 50 to 70 nm, 60 to 70 nm, 30 to 60 nm, 40 to 60 nm, or 50 to 60 nm. When the thickness of the magnetic shell layer (420) exceeds 70 nm, the size of the nanoparticles (P) arranged inside the structure (400) may become excessively large, and the reproducibility of surface-enhanced Raman scattering may deteriorate. When the thickness is less than 30 nm, the gap between the shell layer and the core (410) may become excessively narrow, and the reproducibility of surface-enhanced Raman scattering may deteriorate.
[0084] Referring to FIG. 4b, in one embodiment, the nanoparticles (P) may be formed on the outside of the magnetic shell layer (420) and may include a polymer coating layer (425) surrounding the magnetic shell layer (420). The polymer coating layer (425) imparts a negative charge to the surface of the nanoparticles (P) and may prevent the nanoparticles (P) from clumping together.
[0085] By having such a structure, in the manufacturing process of the structure (400), the nanoparticles (P) can be prevented from clumping together at a stage before forming the silica shell layer (430), and further, the silica shell layer (430) can be formed more easily.
[0086] In one embodiment, the polymer coating layer (425) is not particularly limited, and any polymer coating layer (425) that can be easily selected by a technician with ordinary knowledge in the art should be interpreted as falling within the scope of the present invention.
[0087] Meanwhile, the polymer coating layer (425) may include a water-soluble polymer compound. Examples of the water-soluble polymer compound include, but are not limited to, polyvinylpyrrolidone, polyvinyl alcohol, polyfluorosulfonate, hydroxyethyl cellulose, hydroxypropyl cellulose, cellulose acetate, and polyamide.
[0088] Referring to FIG. 4c, in one embodiment, the structure (400) may further include a biomolecule material (B) disposed in the receiving space (S). With such a structure, by disposing the biomolecule material (B) together within the silica shell layer (430), the biomolecule material (B) disposed within the receiving space (S) and the nanoparticles (P) can come into contact, and as described below, a surface-enhanced Raman scattering spectroscopy signal can be obtained from the surface of the nanoparticles (P).
[0089] As used herein, “biomolecular substance” means a molecule and / or ion existing in a living organism, and may mean a molecule and / or ion synthesized and / or produced during a general biological process such as cell division, morphogenesis, or development.
[0090] Biomolecular substances (B) may include not only proteins, carbohydrates, lipids, nucleic acids, and minerals, but also low-molecular substances such as primary metabolites, secondary metabolites, natural products, blood, serum, plasma, spinal fluid, urine, tissues, and cells, and may include both endogenous and exogenous substances. Examples of biomolecular substances (B) include, but are not limited to, CA 19-9, amyloid beta40, and amyloid beta42 in plasma.
[0091] Meanwhile, as in the structure described above, the structure (400) according to one embodiment may include nanoparticles (P) that non-specifically cause a plasmon resonance phenomenon on the surface with biomolecule materials (B) to form a peak. In an optional embodiment, referring to FIG. 4d, different biomolecule materials (B) may be provided inside the structure (400). Through such a structure, multiple surface-enhanced Raman scattering spectroscopy signals may be obtained from one structure (400).
[0092] Meanwhile, the acquired multiple surface-enhanced Raman scattering spectroscopy signals can be made into a database (DB) through multiple measurements, and based on the databased signals, specific peaks of diseases can be measured, and further, when used in diagnostic kits, etc., it can provide the effect of accurately diagnosing diseases in a short period of time.
[0093] In one embodiment, the structure (400) may include two or more types of biomolecular substances (B). By separating and extracting the upper or lower layer region of the biosolution treated with the pretreatment solution, the dimension of the total surface-enhanced Raman scattering spectroscopy signal obtained can be varied, thereby enabling a more detailed database of the biomolecular substances (B). Ultimately, a structure (400) capable of obtaining a more accurate surface-enhanced Raman scattering spectroscopy signal can be provided.
[0094] As described above, referring again to FIGS. 4A to 4D , in one embodiment, the structure (400) may include a silica shell layer (430). The silica shell layer (430) has an accommodation space (S) therein, and nanoparticles (P) and biomolecule materials (B) may be placed in an area within the accommodation space (S) inside the silica shell layer (430).
[0095] When nanoscale particles are dispersed in a given material, they may clump together due to interactions (e.g., attractive forces) between the particles within the composition. However, in the case of a structure such as this, in which a silica shell layer (430) surrounding individual nanoparticles (P) is provided, the silica shell layer (430) can prevent interactions between the nanoparticles (P), thereby preventing agglomeration of the particles. Furthermore, it can provide an effect of stabilizing the measurement results by enhancing the surface-enhanced Raman scattering spectroscopy signal or increasing the reproducibility of the surface-enhanced Raman scattering spectroscopy signal.
[0096] In addition, the structure (400) according to one embodiment does not necessarily require a substrate, location, or probe that specifically binds to biomolecular substances (B), and thus, when the biomolecular substances (B) are positioned or come into contact with the surface of the nanoparticles (P), the surface-enhanced Raman scattering spectroscopy signals of a significant number of biomolecular substances (B) disposed within the structure (400) can be measured.
[0097] As described above, the measured surface-enhanced Raman scattering spectroscopy signals can be databased, ultimately providing a diagnostic kit for diagnosing diseases even without a biomarker.
[0098] In one embodiment, the silica shell layer (430) may refer to a shell layer formed mainly with silicon (Si) and oxygen (O), and the material thereof is not particularly limited, and any silica shell layer (430) that can be easily selected by a technician with ordinary knowledge in the art should be interpreted as falling within the scope of the present invention. For example, there is a silica shell layer (430) formed by a polymerization reaction of TEOS, but is not limited thereto.
[0099] In one embodiment, the silica shell layer (430) may have a thickness of 20 to 30 nm, and as another optional embodiment, 20 to 29, 20 to 28, 20 to 27, 20 to 26, 20 to 25, 20 to 24, 20.5 to 30, 20.5 to 29, 20.5 to 28, 20.5 to 27, 20.5 to 26, 20.5 to 25, 20.5 to 24, 21 to 30, 21 to 29, 21 to 28, 21 to 27, 21 to 26, 21 to 25, 21 to 24, 21.5 to 30, 21.5 to 29, 21.5 to 28, 21.5 to 27, 21.5 to 26, 21.5 to 25, 21.5 to 24, 22 to 30, 22 to 29, 22 to 28, 22 to 27, 22 to 26, 22 to 25, 22 to 24, 22.5 to 30, 22.5 to 29, 22.5 to 28, 22.5 to 27, 22.5 to 26, 22.5 to 25, 22.5 to 24, 23 to 30, 23 to 29, 23 to 28, 23 to 27, 23 to 26, 23 to 25, 23 to 24, 23.5 to 30, 23.5 to 29, It may have 23.5 to 28, 23.5 to 27, 23.5 to 26, 23.5 to 25, or 23.5 to 24.
[0100] Additionally, in an optional embodiment, the thickness of the silica shell layer (430) may be 23.28 nm or 21.85 nm, and this thickness may vary depending on whether or not ethanol pretreatment is performed.
[0101]
[0102] FIG. 5 is a drawing for explaining an example of a composition for surface-enhanced Raman scattering spectroscopy according to one embodiment.
[0103] Referring to FIG. 5, another embodiment of the present invention can provide a composition (500) for surface-enhanced Raman scattering spectroscopy (hereinafter referred to as a “detection composition”) including a structure (400) for surface-enhanced Raman scattering spectroscopy (hereinafter referred to as a “structure”) dispersed within the composition (500). Here, the composition (500) for detection can be in the form of a liquid or solid containing the structure (400) for surface-enhanced Raman scattering spectroscopy.
[0104] In one embodiment, a structure (400) may be dispersed within the composition for detection (500), and a more specific and detailed description of the structure (400) is replaced with the description in the preceding embodiments.
[0105] In one embodiment, the detection composition (500) may refer to a liquid or a substance in which nanoscale particles can be dispersed, but is not limited thereto. Any liquid or solid substance capable of dispersing nanoparticles (P) that can be easily selected by a person skilled in the art should be interpreted as falling within the scope of the present invention. Meanwhile, the detection composition (500) in a liquid state may be water (H2O), ethanol (EtOH), etc., but is not limited thereto.
[0106] In one embodiment, the detection composition (500) may further include an additive. The additive may be, for example, a surfactant. The surfactant may allow the nanoscale particles and / or the aforementioned structures (400) to be more uniformly dispersed within the detection composition (500).
[0107] In one embodiment, the surfactant is not particularly limited, and cationic surfactants and / or anionic surfactants may be used. In addition, it should be understood that any surfactant that can be easily selected by a person skilled in the art is within the scope of the present invention. Meanwhile, surfactants include, but are not limited to, CTAB, CTAC, etc.
[0108] In addition, in an optional embodiment, a method for manufacturing a structure can proceed with the process of growing nanoparticles in a solution and forming a silica shell layer surrounding the nanoparticles, thereby providing a method for manufacturing a composition (500) for surface-enhanced Raman scattering spectroscopy.
[0109] In one embodiment, the step of forming a solution in which nanoparticles are dispersed may include the step of forming a mixed solution by mixing a first solution containing a first precursor compound including a first plasmonic metal and a second solution containing a second precursor compound including a second plasmonic metal.
[0110] By mixing the first precursor compound and the second precursor compound, the first plasmonic metal of the first precursor compound can grow into a core, and the second plasmonic metal of the second precursor compound can grow into a magnetic shell layer along the outer surface of the core. This allows for the formation of core-shell nanoparticles.
[0111] A more specific and detailed description of the first plasmonic metal and the second plasmonic metal is replaced by the description in the preceding examples.
[0112] Meanwhile, the step of forming core-shell nanoparticles is not particularly limited, and any method that can be easily selected by a technician with ordinary knowledge in the art should be interpreted as falling within the scope of the present invention.
[0113] In one embodiment, the first precursor compound may refer to a compound that includes a first plasmonic metal and can form nanoscale nanoparticles including the first plasmonic metal. This is not particularly limited, and it should be interpreted that any precursor compound of a nanoparticle that can be easily selected by a person skilled in the art is within the scope of the present invention. Examples of the first precursor compound include, but are not limited to, HAuCl4·3H2O.
[0114] In one embodiment, the second precursor compound may refer to a compound that includes a second plasmonic metal and can form nanoscale nanoparticles or shells of nanoparticles including the second plasmonic metal. This is not particularly limited, and any precursor compound of a nanoparticle that can be easily selected by a person skilled in the art should be interpreted as falling within the scope of the present invention. Examples of the second precursor compound include, but are not limited to, AgNO3.
[0115] In one embodiment, when mixing the first solution and the second solution, at least one of a reducing agent and a surfactant may be mixed together. When the mixed solution is formed, the first precursor compound forms a core of a nanoparticle through nanogrowth. When at least one of a reducing agent and a surfactant is mixed together, the growth of the nanoparticle can be further activated.
[0116] The reducing agent and surfactant are not particularly limited, and it should be interpreted that all reducing agents and surfactants that can be easily selected by a person skilled in the art are within the scope of the present invention, and examples thereof include NaBH4, CTAB, and CTAC, respectively.
[0117] In one embodiment, the method for manufacturing a structure may further include a step of mixing a mixed solution and a coating solution containing a water-soluble polymer compound. When the water-soluble polymer compound is mixed into the mixed solution containing nanoparticles, the water-soluble polymer compound forms a coating layer on at least one region of the surface of the nanoparticles, thereby preventing the nanoparticles from agglomerating with each other.
[0118] In one embodiment, the water-soluble polymer compounds include, but are not limited to, polyvinyl pyrrolidone, polyvinyl alcohol, polyfluorosulfonate, hydroxyethyl cellulose, hydroxypropyl cellulose, cellulose acetate, and polyamide.
[0119] Meanwhile, in order to form a coating layer on the surface of the nanoparticles, as an optional example, a mixed solution may be formed and the coating solution may be mixed after a predetermined period of time.
[0120] In one embodiment, the method for manufacturing the structure may further include a step of mixing a coating solution comprising a polymer compound comprising a unit represented by the following chemical formula 1 into the mixed solution, and as an optional embodiment, may further include a step of mixing a coating solution comprising a water-soluble polymer compound comprising a unit represented by the following chemical formula 1:
[0121] [Correction pursuant to Rule 91, July 31, 2025]
[0122] In one embodiment, a method for manufacturing a structure may include a step of mixing a solution or mixed solution in which nanoparticles are dispersed, a third solution containing a biomolecule material, and a fourth solution containing a silica precursor compound. By mixing and reacting the biomolecule material and the silica precursor compound with the solution containing nanoparticles, the nanoparticles and the biomolecule material can be arranged together within the silica shell layer.
[0123] In one embodiment, the biomolecule material and the silica precursor compound can be mixed and reacted in a solution containing nanoparticles using a stirrer, and as an optional example, the mixture can be mixed and reacted in an orbital shaker or a seesaw shaker.
[0124] Other more specific and detailed descriptions of biomolecular materials are replaced by the descriptions in the preceding examples.
[0125] In one embodiment, the silica precursor compound may refer to a compound capable of forming a silica shell layer after a reaction, and this is not particularly limited, and any silica precursor compound that can be easily selected by a person skilled in the art should be interpreted as falling within the scope of the present invention. Examples of such compounds include, but are not limited to, TEOS.
[0126] In one embodiment, the third solution containing the biomolecule material may be a solution treated with a pretreatment solution. The pretreatment solution may comprise substances with low polarity, and optionally, may comprise substances with a lower polarity than water.
[0127] The third solution can be prepared by treating and / or processing a solution directly extracted from a living organism (hereinafter referred to as a “biological solution”). The biological solution may include, for example, blood, serum, plasma, spinal fluid, cells, tissue fluid, urine, etc.
[0128] These biosolutions can use water (H2O) as a solvent, and when a pretreatment solution containing a substance with low polarity is mixed with the biosolution, biomolecules can be separated according to their molecular weight among biomolecules with different molecular weights in the biosolution.
[0129] In one embodiment, the substance having a lower polarity than water includes, but is not limited to, aprotic water, methanol, ethanol, acetone, DMSO, etc., and any substance having a lower polarity than water that can be easily selected by a person skilled in the art should be interpreted as falling within the scope of the present invention.
[0130] In one embodiment, the third solution may include two or more types of biomolecule materials.
[0131] By separating and extracting the upper or lower regions of a biosolution treated with a pretreatment solution, the dimensions of the obtained overall surface-enhanced Raman scattering spectroscopy signal can be diversified, enabling a more detailed database of biomolecular materials. Ultimately, a structure (400) capable of more accurately diagnosing diseases can be provided.
[0132] For example, the processor (220) can obtain Raman signal data obtained by irradiating light on a material including a Raman signal intensifier, and the material including the Raman signal intensifier can mean a material including the above-described structure and a liquid or solid composition and a biological solution of a subject such as plasma, blood, serum, urine, etc.
[0133] As described above, the metal nanostructure may include a biomolecular material in one region of the interior, and when light is irradiated between the biomolecular material arranged inside the metal nanostructure and one region of the metal nanostructure, a surface plasmon phenomenon occurs, and a specific Raman signal according to the type of biomolecular material may be generated in one region of the metal nanostructure.
[0134] Furthermore, a region within a metal nanostructure may contain different biomolecular substances. These different biomolecular substances generate distinct Raman signals. Furthermore, these Raman signals are unique to each biomolecular substance.
[0135] Accordingly, the processor (220) can obtain Raman signal data from the biological solution of the subject by the method described above with reference to FIGS. 4a to 4d and FIG. 5.
[0136] For example, the processor (220) can obtain multiple Raman signal data or single Raman signal data using one biological solution obtained from a subject.
[0137] Hereinafter, with reference to FIGS. 6 to 8, a method for a processor (220) to obtain disease information of a subject using multiple Raman signal data obtained from one biological solution of the subject will be described, and with reference to FIGS. 9 and 10, a method for a processor (220) to obtain disease information of a subject using single Raman signal data obtained from one biological solution of the subject will be described.
[0138]
[0139] Again, referring to FIG. 3, in step 320, the processor (220) may perform a predetermined preprocessing on a plurality of Raman signal data to generate a plurality of Raman signal derivative data corresponding to each of the plurality of Raman signal data.
[0140] FIG. 6 is a flowchart illustrating an example of a method for generating multiple Raman signal derivative data according to one embodiment.
[0141] Hereinafter, with reference to FIG. 6, an example of a method for a processor (220) to generate multiple Raman signal derivative data will be described.
[0142] Referring to FIG. 6, in step 610, the processor (220) can extract some Raman spectrum data related to a specific disease from among the Raman spectrum data included in each of the plurality of Raman signal data based on the size of the Raman signal.
[0143] For example, the processor (220) can selectively extract meaningful information related to a specific disease from the Raman spectrum data. Here, meaningful information related to a specific disease may refer to specific information corresponding to a specific disease from the Raman spectrum data.
[0144] As an example, the processor (220) may extract signals that appear specifically for each specific disease in units of the intervals in which the signals are concentrated. As another example, the processor (220) may extract signals that appear specifically for each specific disease in units of individual Raman shift values.
[0145] For example, if a subject is suffering from a specific disease, Raman spectrum data may include a Raman signal of a specific size corresponding to the specific disease, and the processor (220) may extract only a portion of the Raman spectrum data including the Raman signal of the specific size as described above from among the Raman spectrum data, and remove the remaining Raman spectrum data. Accordingly, the signal-to-noise ratio of the entire data may be improved by having the processor (220) extract only meaningful information from among the Raman spectrum data.
[0146] Therefore, if a Raman signal of a size corresponding to a specific disease is included in the Raman spectrum data included in the Raman signal data, the processor (220) can extract some of the Raman spectrum data in which the corresponding Raman signal appears.
[0147] In other words, the processor (220) can acquire only Raman spectrum data of a portion of the Raman spectrum data where Raman signals of a size corresponding to a specific disease are concentrated. As a result, the signal-to-noise ratio of the Raman signal data can be improved.
[0148] At step 620, the processor (220) may normalize some of the extracted Raman spectrum data.
[0149] For example, the processor (220) can reduce the deviation between Raman spectrum data by normalizing some of the Raman spectrum data using a normalization method such as mean normalization or standard deviation normalization (Z-score normalization). In addition, the method by which the processor (220) normalizes some of the Raman spectrum data is not limited thereto, and any method such as maximum normalization, area normalization, and vector normalization can be used.
[0150] In step 630, the processor (220) may image some of the normalized Raman spectrum data through a predetermined preprocessing to generate a plurality of Raman signal derived data.
[0151] For example, the processor (220) can generate Raman signal derivative data by converting some of the normalized Raman spectrum data into a 2D image using a transformation technique such as GASF (gramian angular summation field), GADF (gramian angular difference field), or MTF (markov transition field).
[0152] In other words, the processor (220) can generate 2D Raman signal derivative data by imaging 1D Raman spectrum data using a predetermined image conversion technique.
[0153] In addition, the processor (220) can extract some Raman spectrum data related to a specific disease for each of the plurality of Raman signal data, normalize each of the extracted some Raman spectrum data, and then generate a plurality of Raman signal derivative data corresponding to each of the plurality of Raman signal data.
[0154] Accordingly, the processor (220) can use the generated plurality of Raman signal derivative data as input data in the inference process of the diagnostic model.
[0155] Meanwhile, the processor (220) can augment multiple Raman signal derived data so that the diagnostic model can use it as learning data for diagnosing a disease during the learning process of the diagnostic model.
[0156] For example, the processor (220) can enhance a plurality of 2D image-ized Raman signal-derived data using a predetermined image enhancement technique. Here, the predetermined image enhancement technique may include, but is not limited to, techniques such as Affine transform, Gaussian Blur, Random erasing, and Sharpening.
[0157] Therefore, the processor (220) can effectively learn a diagnostic model even with a small amount of data by amplifying multiple Raman signal derived data.
[0158]
[0159] Again, referring to FIG. 3, in step 330, the processor (220) can input a plurality of Raman signal derived data as input data into the diagnostic model and obtain a plurality of disease information vectors corresponding to each of the plurality of Raman signal derived data as output data.
[0160] FIG. 7 is a diagram illustrating an example of a method for obtaining multiple disease information vectors according to one embodiment.
[0161] Hereinafter, with reference to FIG. 7, an example of a method for a processor (220) to obtain multiple disease information vectors will be described.
[0162] Referring to FIG. 7, the processor (220) can input a plurality of Raman signal derived data (711, 721, 731) into the diagnostic model (740) as input data of the diagnostic model (740).
[0163] For example, the processor (220) can obtain a plurality of Raman signal data (710, 720, 730) from one biological solution, and perform a predetermined preprocessing on the obtained plurality of Raman signal data (710, 720, 730) to generate a plurality of Raman signal derived data (711, 721, 731).
[0164] Accordingly, the processor (220) can input the generated plurality of Raman signal derived data (711, 721, 731) as input data to the diagnostic model (740).
[0165] In addition, the processor (220) can calculate a probability value that the subject is normal (i.e., a probability value that the subject does not have a disease) and a probability value that the subject has at least one disease for each of the plurality of input Raman signal derived data (711, 721, 731). Here, the disease may include, but is not limited to, information on degenerative brain diseases such as Alzheimer's disease and Parkinson's disease, cancers such as pancreatic cancer, thyroid cancer, prostate cancer, breast cancer, stomach cancer, colon cancer, and the like, and infectious diseases such as hepatitis, coronavirus, influenza, and the like.
[0166] For example, the processor (220) can obtain a plurality of disease information vectors (712, 722, 732) including a probability value that the subject is normal and a probability value that the subject has at least one disease for each of a plurality of Raman signal derived data (711, 721, 731) as output data of the diagnostic model (740). Here, the plurality of disease information vectors (712, 722, 732) can mean a plurality of confidence vectors.
[0167]
[0168] Referring again to FIG. 3, in step 340, the processor (220) can obtain disease information of the subject based on a plurality of disease information vectors (712, 722, 732).
[0169] FIG. 8 is a diagram illustrating an example of a method for obtaining disease information of a subject based on a plurality of disease information vectors according to one embodiment.
[0170] Hereinafter, with reference to FIG. 8, an example of a method in which a processor (220) obtains disease information of a subject using a plurality of disease information vectors is described.
[0171] Referring to Fig. 8, each of the plurality of disease information vectors (812, 822, 832) may include a first dimension (813, 823, 833), a second dimension (814, 824, 834) to a k-th dimension (k is a natural number greater than or equal to 3) (815, 825, 835). Meanwhile, Fig. 8 is a drawing for explaining a case where at least two diseases are detected, and in this case, k must be a natural number greater than or equal to 3. However, it is obvious to those skilled in the art that k may be 2 when one disease is detected.
[0172] For example, the processor (220) can determine the number of dimensions included in each of the plurality of disease information vectors (812, 822, 832) based on the number of diseases to be detected.
[0173] For example, when the number of diseases to be detected is i (i is a natural number greater than or equal to 1), the processor (220) may determine the number of dimensions included in each of the plurality of disease information vectors (812, 822, 832) as (i+1). In other words, when the number of diseases to be detected is i, the processor (220) may determine the number of dimensions of each of the plurality of disease information vectors (812, 822, 832) to include the probability values for cases where the subject is normal and for having the disease to be detected as (i+1). That is, when the number of diseases to be detected is i (i is a natural number greater than or equal to 1), the number of dimensions k of each of the plurality of disease information vectors (812, 822, 832) may be determined as (i+1).
[0174] Accordingly, the processor (220) can determine the number of dimensions included in each of the plurality of disease information vectors (812, 822, 832) based on the number of diseases to be detected, and each dimension can include a probability value that the subject is normal or a probability value that the subject has the disease to be detected. That is, one of the dimensions included in each of the plurality of disease information vectors (812, 822, 832) can include a probability value that the subject is normal, and each of the remaining dimensions can include a probability value that the subject has one disease.
[0175] For example, when the number of diseases to be detected is 1, the processor (220) may determine the number of dimensions of each of the plurality of disease information vectors (812, 822, 832) to be 2, and one dimension of each of the plurality of disease information vectors (812, 822, 832) may include a probability value that the subject has the corresponding disease, and the remaining dimension may include a probability value that the subject is normal. Accordingly, in this case, k in FIG. 8 may be 2.
[0176] As another example, when the number of diseases to be detected is three, the processor (220) may determine the number of dimensions of each of the plurality of disease information vectors (812, 822, 832) to be four, and three dimensions of each of the plurality of disease information vectors (812, 822, 832) may include the probability that the subject has a specific disease, and the remaining dimension may include the probability that the subject is normal. Accordingly, in this case, k in FIG. 8 may be 3. More specifically, when the number of detected diseases is three, the first dimension (813, 823, 833) of each of the plurality of disease information vectors (812, 822, 832) may include a probability value that the subject has the first disease, the second dimension (814, 824, 834) may include a probability value that the subject has the second disease, the third dimension may include a probability value that the subject has the third disease, and the fourth dimension may include a probability value that the subject is normal.
[0177] For example, the processor (220) can aggregate the probability values that the subject has at least one disease included in each of the plurality of disease information vectors (812, 822, 832) through a predetermined method.
[0178] More specifically, the processor (220) can integrate probability values of the same dimension (i.e., dimensions including probability values for the same disease) included in each of the plurality of disease information vectors (812, 822, 832) in a predetermined manner to produce a final disease information vector. Accordingly, the processor (220) can derive a dimension including the largest probability value among the probability values included in the produced final disease information vector, and obtain the disease name or normality corresponding to the dimension as the disease information of the subject.
[0179] Here, the method by which the processor (220) integrates probability values of the same dimension included in each of the plurality of disease information vectors (812, 822, 832) may include, but is not limited to, a method of calculating the average of probability values, a method of adding together based on softmax confidence, a method of majority voting, etc.
[0180] As an additional example, the processor (220) may integrate the probability values of having a specific disease included in each dimension of each of the plurality of disease information vectors (812, 822, 832) to calculate an average value of the probability that the subject has each of the target diseases, and obtain the disease information of the subject based on the calculated average value. In other words, the processor (220) may not only obtain the specific disease with the highest probability that the subject has had as the final diagnosis result, but may also obtain the probability that the subject has had each of the target diseases.
[0181] For example, the processor (220) can calculate, for each dimension (or each disease), an average value of the probability values that the subject has the first disease included in the first dimension (813, 823, 833) of each of the plurality of disease information vectors (812, 822, 832), an average value of the probability values that the subject has the second disease included in the second dimension (814, 824, 834), or an average value of the probability values that the subject has the k-th disease included in the k-th dimension (815, 825, 835).
[0182] In other words, the processor (220) can obtain the probability that the subject has each of the target diseases based on the average value calculated for at least one disease.
[0183] Accordingly, the processor (220) can obtain disease information including whether the subject is normal or the name of a specific disease that the subject has, and the disease information can include information related to a specific disease, such as the expected time of contracting a specific disease and a treatment method for a specific disease.
[0184] Through the above-described method, even when using the same biological solution, it is possible to obtain highly reliable disease diagnosis results by minimizing errors that may occur due to the concentration of Raman signals in only a specific area of the biological solution.
[0185]
[0186] Hereinafter, with reference to FIGS. 9 and 10, as described above, a method for obtaining disease information of a subject by using a single Raman signal data obtained from a single biological solution of the subject by a processor (220) is described.
[0187] For example, the processor (220) can obtain a single Raman signal data from a biological solution of a subject using the method described above with reference to FIGS. 4a to 4d and FIG. 5.
[0188] For example, when a single Raman signal data is acquired using a single biological solution obtained from a subject, the processor (220) can generate first Raman signal derived data and second Raman signal derived data using the single Raman signal data. As an example, the processor (220) can extract features from the single Raman signal data through a predetermined preprocessing to generate the first Raman signal derived data. As another example, the processor (220) can generate second Raman signal derived data by imaging the single Raman signal data through a predetermined preprocessing.
[0189] In addition, the processor (220) can generate first Raman signal derived data and second Raman signal derived data and use them as input data in the learning process and inference process of the diagnostic model. Meanwhile, the diagnostic model in the present disclosure may mean a single diagnostic model when multiple Raman signal data are acquired using one biological solution obtained from the subject, and at least one of the first diagnostic model and the second diagnostic model when a single Raman signal data is acquired using one biological solution. In addition, the diagnostic model may mean an artificial intelligence model learned to detect a disease of a subject by using Raman signal data of the patient as learning data.
[0190] FIG. 9 is a diagram illustrating an example of a method for learning a diagnostic model for diagnosing a disease according to one embodiment.
[0191] Hereinafter, with reference to FIG. 9, an example of a method in which a processor (220) uses Raman signal derived data as input data in the learning process of a diagnostic model is described.
[0192] Referring to FIG. 9, the diagnostic model may be at least one of a first diagnostic model (910) and a second diagnostic model (920), and the first diagnostic model (910) and the second diagnostic model (920) may be trained using first Raman signal derived data (911) and second Raman signal derived data (921) corresponding to the first diagnostic model (910) and the second diagnostic model (920), respectively, as input data.
[0193] For example, the processor (220) may perform a predetermined preprocessing to use the Raman signal data (900) of the patient as input data in the learning process of the diagnostic model. Here, the Raman signal data of the patient may be data acquired by the method described above with reference to FIGS. 4A to 4D and FIG. 5. In addition, the processor (220) may preprocess the Raman signal data (900) of the patient to generate first Raman signal derived data (911) and second Raman signal derived data (921) in order to use the Raman signal data (900) of the patient as input data in the learning process of the diagnostic model. That is, the processor (220) may use the first Raman signal derived data (911) and second Raman signal derived data (921) generated by preprocessing the Raman signal data (900) of the patient as input data in the learning process of the diagnostic model.
[0194] As an example, the processor (220) may preprocess the Raman signal data (900) of a patient to generate first Raman signal derived data (911) for learning a diagnostic model. Specifically, the processor (220) may augment the Raman signal data (900), extract features of the augmented Raman signal data (900), and then adjust an unbalanced class distribution of the Raman signal data (900) to generate first Raman signal derived data (911).
[0195] First, the processor (220) can enhance the Raman signal data (900) of the patient. For example, the processor (220) can prevent overfitting of the diagnostic model by arbitrarily injecting noise into the Raman signal data (900) of the patient used as input data in the learning process. In addition, the processor (220) can scale the Raman signal data (900), i.e., adjust the size of the Raman signal data (900). For example, the processor (220) may divide the enhanced Raman signal data (900) into x equal parts (x is a natural number greater than or equal to 1) based on at least one of a principal component analysis technique (e.g., PCA (principal component analysis)) and an independent component analysis technique (e.g., ICA (independent component analysis)), and extract y features (y is a natural number greater than or equal to 1) for each of the x equal parts to generate Raman signal derived data (911, 921), but is not limited thereto. In addition, the processor (220) may balance the unbalanced class distribution of the Raman signal data (900) from which features are extracted based on a synthetic minority oversampling technique (SMOTE) technique.
[0196] As another example, the processor (220) may preprocess the Raman signal data (900) of a patient to generate second Raman signal derived data (921) for learning a diagnostic model. Specifically, the processor (220) may augment the Raman signal data (900), image the augmented Raman signal data (900), and then augment the imaged Raman signal data (900) again to generate second Raman signal derived data (921).
[0197] First, the processor (220) can enhance the Raman signal data (900) of the patient. For example, the processor (220) can prevent overfitting of the diagnostic model by arbitrarily injecting noise into the Raman signal data (900) of the patient used as input data in the learning process. In addition, the processor (220) can scale the Raman signal data (900), that is, adjust the size of the Raman signal data (900). For example, the processor (220) can generate second Raman signal derived data (921) by 2D imaging the enhanced Raman signal data (900) based on at least one of a gramian angular summation field (GASF) and a gramian angular difference field (GADF), but is not limited thereto. In addition, the processor (220) can delete some data from the imaged Raman signal data (900) to enhance the Raman signal data (900) again. In other words, the processor (220) can two-dimensionally image the Raman signal data (900) and enhance the Raman signal data (900) by deleting some data from each of the horizontal and vertical axes of the two-dimensional image.
[0198] For example, the processor (220) may input first Raman signal derived data (911) and second Raman signal derived data (921) as input data into the first diagnostic model (910) and the second diagnostic model (920), respectively, to obtain output data (930, 940). Here, the output data (930, 940) may mean the probability that a patient with a disease has a disease.
[0199] Accordingly, the processor (220) can learn the first diagnostic model (910) and the second diagnostic model (920) so as to accurately diagnose the disease possessed by the patient based on the input data and output data of each of the first diagnostic model (910) and the second diagnostic model (920).
[0200] For example, the processor (220) can preprocess Raman signal data of a patient with dementia to generate Raman signal derived data, and input the generated Raman signal derived data as input data into the diagnostic model (910, 920). In addition, the processor (220) can obtain the probability that the patient has dementia as output data (930, 940) from the diagnostic model (910, 920). Accordingly, the processor (220) can learn the diagnostic model (910, 920) so that it can diagnose that the patient has dementia with a high probability. Here, the processor (220) can obtain not only the probability that the patient has dementia, but also the probability that the patient has various diseases such as cancer and infectious diseases as output data.
[0201] In addition, the processor (220) can learn each of the plurality of sub-diagnostic models included in each of the first diagnostic model (910) and the second diagnostic model (920) to be able to diagnose a specific disease with high accuracy. As an example, the processor (220) can learn the 1-1 sub-diagnostic model to be able to accurately diagnose pancreatic cancer. As another example, the processor (220) can learn the 2-1 sub-diagnostic model to be able to accurately diagnose Alzheimer's disease. However, the above-described example is not limited thereto, and the processor (220) can learn each of the plurality of sub-diagnostic models to be able to accurately diagnose various diseases.
[0202]
[0203] FIG. 10 is a drawing for explaining an example of a method for obtaining disease information of a subject according to one embodiment.
[0204] Hereinafter, with reference to FIG. 10, an example of a method in which a processor (220) obtains disease information of a subject using a diagnostic model is described.
[0205] Referring to FIG. 10, the processor (220) can obtain disease information (1060) of a subject as output data (1030, 1040) of a diagnostic model (1010, 1020). Here, the disease information (1060) of the subject may mean the probability that the subject has at least one disease, but is not limited thereto. Accordingly, the processor (220) can obtain the probability that the subject has a specific disease among a plurality of diseases as output data of the diagnostic model (1010, 1020). In addition, the disease may include, but is not limited to, information on degenerative brain diseases such as Alzheimer's disease and Parkinson's disease, cancer diseases such as pancreatic cancer, thyroid cancer, prostate cancer, breast cancer, stomach cancer, and colon cancer, and infectious diseases such as hepatitis, coronavirus, and influenza.
[0206] In addition, the first diagnostic model (1010) and the second diagnostic model (1020) may each include at least one sub-diagnostic model. As an example, the first diagnostic model (1010) may include a first sub-diagnostic model, and the first sub-diagnostic model may include a 1-1 sub-diagnostic model to a 1-m sub-diagnostic model (m is a natural number greater than or equal to 1). As another example, the second diagnostic model (1020) may include a second sub-diagnostic model, and the second sub-diagnostic model may include a 2-1 sub-diagnostic model to a 2-n sub-diagnostic model (n is a natural number greater than or equal to 1).
[0207] For example, the processor (220) can preprocess the Raman signal data of the subject in the above-described manner with reference to FIG. 9 to generate Raman signal derivative data of the subject, and can use the generated Raman signal derivative data of the subject as input data of the first diagnostic model (1010), the second diagnostic model (1020), or the first sub-diagnostic model or the second sub-diagnostic model included in each of the first diagnostic model (1010) and the second diagnostic model (1020).
[0208] In addition, the processor (220) may generate Raman signal derivative data of a subject by performing at least one operation among a plurality of operations for preprocessing the Raman signal data described above with reference to FIG. 9. Accordingly, the processor (220) may input the Raman signal derivative data of the subject as input data into a diagnostic model and obtain disease information (1060) of the subject as output data. Here, the Raman signal data of the subject may be data obtained by the method described above with reference to FIGS. 4A to 4D and FIG. 5.
[0209] For example, the processor (220) can preprocess Raman signal data in different ways to generate first Raman signal derived data and second Raman signal derived data corresponding to the first diagnostic model (1010) and the second diagnostic model (1020), respectively.
[0210] As an example, the processor (220) may extract features from Raman signal data through a predetermined preprocessing to generate first Raman signal derived data. Specifically, the processor (220) may divide the enhanced Raman signal data into x equal parts (x is a natural number greater than or equal to 1) based on at least one of a principal component analysis technique (e.g., principal component analysis (PCA)) and an independent component analysis technique (e.g., independent component analysis (ICA)), and extract y features (y is a natural number greater than or equal to 1) for each of the x equal parts to generate first Raman signal derived data, but is not limited thereto.
[0211] As another example, the processor (220) may generate second Raman signal derived data by imaging the Raman signal data through a predetermined preprocessing. Specifically, the processor (220) may generate second Raman signal derived data by 2D imaging the enhanced Raman signal data based on at least one of a gramian angular summation field (GASF) and a gramian angular difference field (GADF), but is not limited thereto.
[0212] Accordingly, the processor (220) can obtain disease information (1060) by inputting the first Raman signal derived data and the second Raman signal derived data of the subject generated as input data into the first diagnostic model (1010) and the second diagnostic model (1020), respectively.
[0213] For example, the processor (220) can obtain output data (1030, 1040) of each of the first diagnostic model (1010) and the second diagnostic model (1020).
[0214] For example, the processor (220) may input either the first Raman signal derived data or the second Raman signal derived data as input data to at least one of the first sub-diagnostic model and the second sub-diagnostic model included in the diagnostic model (1010, 1020). In other words, the processor (220) may input the first Raman signal derived data to the first sub-diagnostic model, and input the second Raman signal derived data to the second sub-diagnostic model.
[0215] Additionally, the processor (220) can obtain disease information of the subject by combining the output data (1030, 1040) of each of the first sub-diagnosis model and the second sub-diagnosis model.
[0216] As an example, a user can obtain final disease information (1060) by calculating the average and median values of the output data (1030, 1040), i.e., the probability of having a disease. As another example, the processor (220) can obtain disease information (1060) by combining the output data (1030, 1040) using an ensemble model (1050). Here, the ensemble model (1050) may mean a single artificial intelligence model created by combining at least one or more learned artificial intelligence models.
[0217] For example, the processor (220) may input the output data (1030, 1040) of each of the first sub-diagnosis model and the second sub-diagnosis model as input data to the learned ensemble model (1050), and obtain the subject's disease information (1060) as the output data of the ensemble model (1050). As an example, the ensemble model (1050) may be an artificial intelligence model generated by combining the first sub-diagnosis model and the second sub-diagnosis model. As another example, the ensemble model (1050) may be an artificial intelligence model generated by combining at least one artificial intelligence model that has been learned separately.
[0218] For example, the processor (220) can obtain disease information (1060) of the subject using various ensemble techniques.
[0219] As an example, the processor (220) may obtain the subject's disease information (1060) based on a hard voting technique. Here, the hard voting technique may be a technique of outputting the output data with the highest frequency among the output data of each of at least one artificial intelligence model constituting the ensemble model (1050) as the output data of the ensemble model (1050), i.e., the subject's disease information (1060).
[0220] As another example, the processor (220) may obtain the disease information (1060) of the subject based on a soft voting technique. Here, the soft voting technique may be a technique that calculates a weighted average value of the output data of each of at least one artificial intelligence model constituting the ensemble model (1050) and outputs it as the output data of the ensemble model (1050), i.e., the disease information (1060) of the subject. In addition, the weighted average value may refer to an average value calculated by assigning weights to values whose average values are to be calculated. Accordingly, the processor (220) may assign weights to the output data of each of at least one artificial intelligence model constituting the ensemble model (1050) and calculate the average value.
[0221] For example, the processor (220) may assign the same weight to the output data of each of at least one artificial intelligence model constituting the ensemble model (1050). In addition, the processor (220) may assign a weight to the output data of each of the artificial intelligence models constituting the ensemble model (1050) based on the reliability value of each of the artificial intelligence models constituting the ensemble model (1050). In addition, the processor (220) may assign a weight to the output data of each of the artificial intelligence models constituting the ensemble model (1050) using a weight assignment model. Here, the weight assignment model may be a separate artificial intelligence model that is trained to assign a weight based on the characteristics of each of the artificial intelligence models constituting the ensemble model (1050).
[0222] In addition, as described above with reference to FIG. 9, a specific artificial intelligence model can be trained to accurately diagnose a specific disease, and the processor (220) can construct an ensemble model (1050) differently for each disease using at least one artificial intelligence model.
[0223] Accordingly, the processor (220) can obtain disease information of the subject by combining the output data of each of the first sub-diagnosis model and the second sub-diagnosis model.
[0224] In addition, the processor (220) can retrain a diagnostic model (e.g., a first diagnostic model and / or a second diagnostic model) and a sub-diagnostic model (e.g., a first sub-diagnostic model and / or a second sub-diagnostic model) using the acquired disease information and clinical information of the subject as learning data. Here, the clinical information may include, but is not limited to, disease stage information, immune test information of the subject, PET (positron emission tomography) test results, CT (computed tomography) test information, MRI (magnetic resonance imaging) test information, and tissue test information.
[0225] Meanwhile, the above-described method can be written as a program that can be executed on a computer, and can be implemented on a general-purpose digital computer that runs the program using a computer-readable recording medium. In addition, the structure of the data used in the above-described method can be recorded on a computer-readable recording medium through various means. The computer-readable recording medium includes storage media such as magnetic storage media (e.g., ROM, RAM, USB, floppy disk, hard disk, etc.) and optical reading media (e.g., CD-ROM, DVD, etc.).
[0226] Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from the essential characteristics of the above-described invention. Therefore, the disclosed methods should be considered illustrative rather than restrictive. The scope of the claims, not the foregoing description, is defined by the scope of the patent, and should be interpreted to encompass all differences within the scope equivalent thereto.
Claims
1. A step of acquiring multiple Raman signal data from a biological solution of a subject using a Raman signal intensifier; A step of performing a predetermined preprocessing on the plurality of Raman signal data to generate a plurality of Raman signal derivative data corresponding to each of the plurality of Raman signal data; A step of inputting the plurality of Raman signal derived data into a diagnostic model as input data and obtaining a plurality of disease information vectors corresponding to each of the plurality of Raman signal derived data as output data; and A step of obtaining disease information of the subject based on the plurality of disease information vectors; including; A method for diagnosing diseases using artificial intelligence models.
2. In paragraph 1, The above multiple Raman signal data are, A method for obtaining data by irradiating light on a material containing the above Raman signal intensifier.
3. In paragraph 1, The step of generating the above plurality of Raman signal derivative data is: A step of extracting some Raman spectrum data related to a specific disease from among the Raman spectrum data included in each of the plurality of Raman signal data based on the magnitude of the Raman signal; a step of normalizing some of the Raman spectrum data; and A method comprising: a step of generating a plurality of Raman signal derivative data by imaging some of the normalized Raman spectrum data through a predetermined preprocessing; 4. In paragraph 1, The step of obtaining the above multiple disease information vectors is: A step of inputting the plurality of Raman signal derived data into the diagnostic model as the input data of the diagnostic model; A step of calculating a probability value that the subject has at least one disease for each of the plurality of Raman signal derived data; and A method comprising: a step of obtaining the plurality of disease information vectors including the probability values as the output data of the diagnostic model; 5. In paragraph 1, The step of obtaining the above-mentioned subject's disease information is: A step of calculating a final disease information vector by integrating the probability values that the subject has at least one disease included in each of the plurality of disease information vectors; and A method comprising: a step of obtaining disease information of the subject based on the final disease information vector.
6. A computer-readable recording medium recording a program for executing the method of paragraph 1 on a computer.
7. Memory in which at least one program is stored; and comprising at least one processor executing at least one program; At least one processor, A method for obtaining a plurality of Raman signal data from a biological solution of a subject using a Raman signal intensifier, performing a predetermined preprocessing on the plurality of Raman signal data to generate a plurality of Raman signal derived data corresponding to each of the plurality of Raman signal data, inputting the plurality of Raman signal derived data into a diagnostic model as input data, obtaining a plurality of disease information vectors corresponding to each of the plurality of Raman signal derived data as output data, and obtaining disease information of the subject based on the plurality of disease information vectors. A device that diagnoses diseases using artificial intelligence models.
Citation Information
Patent Citations
Method for analyzing biological specimens by spectral imaging
KR1020130056886A
Method for manufacturing contanier and apparatus for manufacturing contanier
KR1020220109368A
Chair with cool and hot air supply
KR1020240057119A
Display having curved shape and electronic device including same
KR1020250026087A
KR20230115631A