Plasma digital marker mining and disease classification device and related equipment
By calculating the influence score of plasma sample spectra and binding to the spectral absorption peak of pathogenic proteins, an iteratively optimized disease classification device was constructed, which solved the problem of unstable and inspecific screening of markers for plasma spectroscopy data in the prior art, and achieved accurate diagnosis of MCI and AD patients, providing a simple and non-invasive early diagnosis tool.
Patent Information
- Application Number
- CN202510619654.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-26
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art is difficult to effectively screen stable, specific and biologically significant digital markers from plasma spectral data for accurate identification of patients with mild cognitive impairment (MCI) and early Alzheimer's disease (AD). Traditional methods have problems with the lack of stability, specificity and biological significance of markers caused by relying on statistical analysis.
By calculating the influence score of plasma sample spectrum, a digital marker sorting cluster was screened, and the spectral absorption peak of pathogenic proteins was combined to construct an iteratively optimized disease classification device. The spectral information of plasma samples was obtained using ATR-FTIR technology, and iterative training was combined with machine learning algorithms to screen out biologically significant plasma digital markers.
It realizes the mining of stable, specific and biologically significant digital markers from plasma spectral data, and constructs a disease classification device that can accurately distinguish MCI and AD patients, providing a simple, non-invasive and economical diagnostic method, which improves the accuracy and interpretability of the diagnosis.
Smart Images

Figure CN120544844A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent medical detection, and specifically to a plasma digital marker mining and disease classification device, a disease screening device, an electronic device and a computer storage medium. Background Art
[0002] Early diagnosis and screening are crucial for intervention in neurodegenerative diseases, especially for identifying patients with mild cognitive impairment (MCI). MCI is considered a precursor to Alzheimer's disease (AD) and other types of dementia. Promptly identifying patients with MCI can provide an opportunity for early intervention and slow disease progression in potential AD patients.
[0003] Currently, the diagnosis of MCI and early AD relies primarily on neuropsychological testing, cerebrospinal fluid biomarker detection, and brain imaging. However, these methods have limitations: neuropsychological test results may be affected by factors such as educational level; cerebrospinal fluid examinations are invasive; and brain imaging is expensive and unsuitable for large-scale screening. Therefore, there is an urgent need to develop a simple, noninvasive, and cost-effective method to identify patients with MCI.
[0004] Blood, as an easily accessible biological sample, contains a wealth of biological information. In recent years, researchers have attempted to use plasma samples combined with spectral analysis technology for disease diagnosis. Attenuated Total Reflectance Fourier Transform Infrared Spectroscopy (ATR-FTIR) technology can obtain the spectral fingerprint of plasma samples, reflecting the comprehensive information of multiple biomolecules. However, how to effectively screen out plasma digital markers that can accurately identify MCI, AD, and non-AD patients from complex plasma spectral data remains a technical problem that needs to be solved urgently.
[0005] Existing methods for screening plasma digital markers primarily rely on statistical methods, such as t-tests and analysis of variance, to compare the differences in spectral intensity at various wavenumbers between samples from different groups and select wavenumbers with significant differences as markers. However, these methods fail to fully consider the interactions and redundant information between markers, and the selected markers may lack biological significance and be difficult to guide clinical practice. Furthermore, markers selected solely by statistical methods may be overly sensitive and difficult to reproduce across different sample sets.
[0006] Therefore, how to extract stable, specific and biologically meaningful plasma digital markers from plasma spectral data and then build a disease classification device that can accurately distinguish MCI patients is a technical problem that needs to be solved urgently. Summary of the Invention
[0007] This application provides a plasma digital marker mining and disease classification device, disease screening device, equipment and medium, which can mine stable, specific and biologically significant digital markers from plasma spectral data, and then construct a disease classification device that can accurately distinguish MCI patients.
[0008] In a first aspect of the present application, the present application provides a plasma digital marker mining and disease classification device, characterized by comprising: a plasma sample spectrum acquisition module, configured to collect spectra of multiple plasma samples from patients with MCI, AD, non-AD, and HC, and calculate an influence score for the spectrum of each plasma sample, and determine the spectrum of the plasma sample with the highest influence score as the digital marker ranking cluster, wherein the influence score represents the similarity of the plasma sample spectrum to the spectra of the plasma samples of MCI, AD, non-AD, and HC; a plasma candidate digital marker selection module, configured to iterate the plasma spectrum samples in the initial disease classification device based on the plasma full spectrum digital marker ranking cluster to obtain a plasma candidate digital marker cluster; A pathogenic protein digital marker cluster acquisition module is used to acquire spectral markers corresponding to spectral absorption peaks of multiple pathogenic proteins related to the MCI, AD and non-AD patients to obtain a pathogenic protein digital marker cluster; A plasma digital marker mining module, configured to determine plasma digital markers for disease classification based on a digital logical combination of the plasma candidate digital marker cluster and the pathogenic protein digital marker cluster; The disease classification device construction module is used to train the iterated initial disease classification device based on the plasma digital markers to obtain a trained disease classification device.
[0009] By adopting the above technical solution, the problem of lack of stability, specificity and biological significance of markers caused by relying solely on statistical analysis in the existing technology is effectively solved, and the goal of extracting stable, specific and biologically significant digital markers from complex plasma spectral data is achieved.
[0010] First, the plasma sample spectrum acquisition module calculates the influence scores of plasma sample spectra and identifies the sample spectra with the highest influence scores as a digital marker ranking cluster, overcoming the limitations of traditional statistical methods. This screening method based on sample spectral feature similarity more comprehensively considers the interactions between markers and statistically selects spectral markers that contribute significantly to the differentiation of MCI, AD, non-AD patients, and healthy controls (HCs).
[0011] Secondly, the plasma candidate digital marker selection module iterates the plasma spectral samples from the initial disease classification device to generate clusters of plasma candidate digital markers, further improving the stability and repeatability of the markers. This addresses the existing problem of markers that rely solely on statistical methods to screen and are difficult to reproduce across different sample sets.
[0012] Furthermore, this application also introduces a pathogenic protein digital marker cluster acquisition module. By obtaining spectral markers corresponding to the spectral absorption peaks of multiple pathogenic proteins related to MCI, AD and non-AD patients, a biologically significant pathogenic protein digital marker cluster is obtained, which makes up for the deficiency that relying solely on statistical methods may result in the screened markers lacking biological significance.
[0013] The plasma digital marker mining module further combines candidate plasma digital marker clusters with pathogenic protein digital marker clusters through digital logic to identify plasma digital markers for disease classification. This module considers both the statistical properties of the markers and prior knowledge of disease-related proteins, resulting in the final selected plasma digital markers possessing stability, specificity, and biological interpretability, effectively addressing the lack of comprehensive characteristics in existing marker technologies.
[0014] Finally, the disease classification device construction module uses the screened plasma digital markers to train the initial disease classification device, resulting in a classification device that can accurately distinguish between MCI, AD, non-AD patients, and HC. This device overcomes the limitations of existing technologies that rely on invasive examinations or high-cost equipment, providing a simple and economical method to identify MCI patients and a powerful tool for early intervention in AD.
[0015] Optionally, the acquisition of spectra of multiple plasma samples from patients with MCI, AD, non-AD patients, and HC includes: Multiple plasma samples were collected from patients with MCI, AD, non-AD and HC, and the plasma samples were analyzed at 4000-600 cm -1 Within the wavenumber range, cumulative spectral scans are performed a preset number of times at a preset resolution to obtain spectra of the multiple plasma samples of the MCI, AD, non-AD patients, and HC.
[0016] By adopting the above technical scheme, multiple plasma samples of MCI, AD, non-AD patients and HC were collected and analyzed at 4000-600 cm -1 Spectral scanning within the wavenumber range can fully obtain the vibration information of various biomolecules in the plasma sample. -1 The wavenumber range covers the mid-infrared spectral region, which is optimal for studying the structure and vibrational modes of organic molecules. It can capture the characteristic absorption peaks of key disease-related biomolecules such as proteins, lipids, and nucleic acids. Scanning at a preset resolution ensures sufficiently detailed spectral data to accurately identify and distinguish subtle spectral differences in different disease states. Furthermore, by accumulating spectral scans over a preset number of times, the signal-to-noise ratio is significantly improved, the impact of random noise is reduced, and the reliability and reproducibility of spectral data are enhanced.
[0017] Optionally, calculating the influence score of the spectrum of each plasma sample and determining the spectra of the plasma sample with the highest influence score as the digital marker ranking cluster includes: For a spectrum of any plasma sample in the spectral information space, respectively calculating the similarity between the spectrum of the plasma sample and spectra of plasma samples of the same type, and the difference between the spectrum of the plasma sample and spectra of plasma samples of different types; The sum of the similarity and difference of the spectra of the plasma samples is used as an influence score, and the spectra of the multiple plasma samples ranked higher in influence scores are determined as a plasma full-spectrum digital marker ranking cluster.
[0018] By adopting the above technical solution, the influence score of each plasma sample spectrum in the spectral information space is calculated, innovatively combining the evaluation of intra-sample similarity and inter-sample difference, thereby effectively identifying the spectral features with the most diagnostic value. Specifically, for each plasma sample spectrum, its similarity with samples of the same type and its difference with samples of different types are calculated simultaneously. This dual consideration ensures that the selected markers not only perform consistently across samples of the same type, but can also effectively distinguish between different disease states. Using the sum of similarity and difference as the influence score further balances the stability and discriminatory power of the feature, avoiding the bias that may be caused by a single indicator.
[0019] Optionally, the step of obtaining spectral markers corresponding to spectral absorption peaks of multiple pathogenic proteins associated with the MCI, AD, and non-AD patients includes: The spectral absorption peak positions of multiple pathogenic protein molecular structures related to the MCI, AD and non-AD patients are determined as digital markers, wherein the protein molecular structure is at least one of a chemical bond and a functional group, and the peak position of the pathogenic protein digital marker cluster is negatively correlated with the biomarker levels of the MCI, AD and non-AD patients.
[0020] By employing this technical approach, the positions of spectral absorption peaks depicting the molecular structures of multiple pathogenic proteins associated with MCI, AD, and non-AD patients were identified as digital biomarkers, cleverly linking spectral analysis directly with pathological changes at the protein molecular level. This not only enhances the biological significance of the digital biomarkers but also improves the interpretability of diagnostic results. Specifically, by focusing on the chemical bonds and functional groups of proteins, subtle changes in protein structure and function, such as misfolding or abnormal aggregation, that occur during the disease process are captured. These changes are directly reflected in spectral signatures at specific wavenumber positions. Importantly, the peak positions of the pathogenic protein digital biomarker clusters were found to be negatively correlated with biomarker levels in MCI, AD, and non-AD patients, revealing an intrinsic link between spectral changes and disease progression. This negative correlation may reflect conformational changes or degradation processes of certain key proteins in plasma as the disease progresses, providing a molecular-level mechanistic explanation for spectroscopic diagnosis. This approach not only improves diagnostic accuracy but also provides a new tool for monitoring disease progression.
[0021] Optionally, the multiple pathogenic proteins include Aβ42, p-tau and GFAP, and the preset wavenumber range includes 1800-1700 cm -1 、1750-1200cm -1 and 1200-900cm -1 .
[0022] By adopting the above technical scheme, Aβ42, p-tau and GFAP were selected as key pathogenic proteins and focused on 1800-1700 cm -1 、1750-1200cm -1 and 1200-900cm -1 These three specific wavenumber ranges effectively combine the core pathological mechanism of AD with ATR-FTIR spectroscopy analysis. Aβ42, p-tau, and GFAP represent amyloid plaque formation, neurofibrillary tangles, and neuroinflammatory responses in AD pathology, respectively. These three processes together constitute the key links in the onset of AD. By analyzing the spectral characteristics of these proteins within a specific wavenumber range, we can fully capture the multi-dimensional changes in the pathological process of AD. Specifically, 1800-1700 cm -1This region mainly reflects the stretching vibration of lipid C=O, which may reveal AD-related lipid metabolism abnormalities; 1750-1200 cm -1 The range includes the amide I and II bands of proteins, especially 1600-1700 cm -1 The amide I band and 1500-1560 cm -1 The amide II band is closely related to the secondary structure of proteins, especially the β-folded structure of Aβ fibers; 1200-900 cm -1 The specific wavenumber ranges of the spectral region may reflect information related to membrane damage and nucleic acid oxidative damage. This precise selection of wavenumber ranges not only improves the specificity of spectral analysis, but also significantly reduces the amount of data required for processing, improving computational efficiency. By focusing on these specific wavenumber ranges, the scheme can simultaneously obtain multiple aspects of information related to AD pathology, such as protein conformational changes, abnormal lipid metabolism, and oxidative stress, providing a multi-dimensional spectral fingerprint for early diagnosis and disease monitoring of AD.
[0023] Optionally, the plasma digital marker mining and disease classification device further comprises: a disease classification model device evaluation module, configured to evaluate the performance of the disease classification device based on the output results of the disease classification device; The method of iterating the plasma spectrum samples in the initial disease classification device based on the plasma full spectrum digital marker ranking cluster to obtain the plasma candidate digital marker cluster includes: According to the influence scores of the digital markers in the plasma full spectrum digital marker ranking cluster, the digital markers are sequentially input into the initial disease classification device, and the impact of each digital marker on the initial disease classification device is evaluated by the disease classification model device evaluation module to obtain an evaluation result; Screening out no less than 20% of the digital markers in the plasma candidate digital marker cluster based on the evaluation results; The iterative initial disease classification device is trained based on the plasma digital markers to obtain a trained disease classification device, including: The plasma digital markers are divided into a training set, a validation set, and a test set. The iterated initial disease classification device is trained using the training set. The parameters of the trained initial disease classification device are adjusted using the validation set. The performance of the trained initial disease classification device is evaluated using the test set and the disease classification model device evaluation module to obtain a trained disease classification device.
[0024] By implementing the aforementioned technical solution, this approach significantly improves the performance and reliability of plasma digital biomarker mining and disease classification by introducing a disease classification model evaluation module and employing an iterative optimization and dynamic evaluation approach. First, digital biomarkers are sequentially fed into the initial disease classification system based on their influence scores within the plasma full-spectrum digital biomarker ranking cluster. The disease classification model evaluation module then evaluates each biomarker's contribution in real time. This dynamic screening approach ensures that the final selected biomarkers are not only statistically significant but also perform well in the actual classification task. By eliminating at least 20% of the digital biomarkers based on the evaluation results, the feature dimensionality is effectively reduced, computational efficiency is improved, and potentially redundant information is removed, enhancing the model's generalization capability. This performance-based dynamic adjustment ensures the scientific and effective nature of the screening process, resulting in a more streamlined and efficient final cluster of candidate plasma digital biomarkers. Second, the strategy of partitioning the plasma digital biomarkers into training, validation, and test sets fully utilizes the data's value and effectively prevents overfitting. The complete process of model training on the training set, parameter adjustment on the validation set, and performance evaluation on the test set ensures good generalization performance of the disease classification system on unseen data.
[0025] Optionally, the performance of the disease classification device includes at least one of specificity, sensitivity, and area of the ROC curve; The specificity and sensitivity of the initial disease classification device after iteration are not less than 80%, and the area under the ROC curve is greater than 75%; The specificity and sensitivity of the trained initial disease classification device are not less than 85%, and the area of the ROC curve approaches 1.
[0026] By employing the above technical solution, the system effectively ensures high performance of the plasma digital biomarker mining and disease classification device at different stages by setting strict performance metrics, including specificity, sensitivity, and ROC curve area. First, the initial disease classification device after iteration is required to have specificity and sensitivity of at least 80%, and an ROC curve area greater than 75%. This standard ensures high diagnostic accuracy from the outset. 80% specificity and sensitivity indicate that the device can effectively identify healthy individuals and detect diseased individuals, significantly reducing the risk of misdiagnosis and missed diagnoses. An ROC curve area greater than 75% indicates that the device maintains good classification performance at different decision thresholds, providing flexibility for clinical application. Furthermore, the initial disease classification device after training is required to have specificity and sensitivity of at least 85%, and an ROC curve area close to 1. This stringent standard ensures that the final model has excellent diagnostic capabilities, capable of distinguishing healthy individuals from patients with neurodegenerative diseases with extremely high accuracy. Specificity and sensitivity exceeding 85% indicate further reductions in misdiagnosis and missed diagnosis rates, while the ROC curve area approaching 1 indicates that the model maintains near-perfect classification performance across a wide range of decision thresholds. This significant performance improvement not only enhances diagnostic reliability but also facilitates early screening and intervention. By gradually increasing performance requirements during model iteration and training, this approach achieves a leap from good to excellent performance, fully realizing the potential of plasma ATR-FTIR spectroscopy in the diagnosis of neurodegenerative diseases.
[0027] In a second aspect of the present application, a disease screening device is provided, comprising a disease recognition model constructed by the plasma digital marker mining and disease classification device described in any one of the above items.
[0028] In a third aspect of the present application, a computer storage medium is provided, which stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executed by the disease identification model constructed by the plasma digital marker mining and disease classification device described in any of the above items.
[0029] In the fourth aspect of the present application, an electronic device is provided, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the disease identification model constructed by the plasma digital marker mining and disease classification device described in any of the above items. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a schematic diagram of the structure of a plasma digital marker mining and disease classification device provided in an embodiment of the present application; Figure 2This is a flow chart of a plasma digital marker mining and disease classification device training process provided in an embodiment of the present application; Figure 3 This is a flow chart of a plasma digital marker mining process of a plasma digital marker mining and disease classification device provided in an embodiment of the present application; Figure 4 This is a schematic diagram of the comparison and diagnosis results of ATR-FTIR spectra of AD and HC provided in the examples of the present application; Figure 5 Schematic diagram of the MCI recognition capability of Fourier transform infrared spectroscopy provided in an embodiment of the present application; Figure 6 This is a schematic diagram of the differential diagnosis performance of Fourier transform infrared spectroscopy provided in an embodiment of the present application; Figure 7 This is a schematic diagram of the linear relationship between the plasma GFAP, p-Tau and Aβ42 levels of AD patients in cohort 1 and cohort 3 and the absorbance of the corresponding AD spectral digital biomarkers provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0032] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0033] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0034] Please refer to Figure 1 , Figure 1This is a schematic diagram of the structure of a plasma digital marker mining and disease classification device provided in an embodiment of the present application. Specifically, the plasma digital marker mining and disease classification device may include: The plasma sample spectrum acquisition module is used to collect spectra of multiple plasma samples from patients with MCI, AD, non-AD patients, and HC, and calculate the influence score of the spectrum of each plasma sample. The spectra of the plasma samples with the highest influence score are determined as the digital marker ranking cluster. The influence score represents the similarity of the plasma sample spectrum to the spectra of plasma samples of MCI, AD, non-AD patients, and HC.
[0035] Specifically, to accurately diagnose Alzheimer's disease (AD) and its early stage, mild cognitive impairment (MCI), this paper proposes a digital biomarker mining method based on spectral analysis of plasma samples. The core of this method lies in the plasma sample spectrum acquisition module, which not only collects spectra from multiple plasma samples of patients with MCI, AD, non-AD patients, and healthy controls (HC), but also selects the most diagnostically valuable spectral features by calculating influence scores.
[0036] Specifically, the plasma sample spectrum acquisition module first scans the collected plasma sample using attenuated total reflectance Fourier transform infrared spectroscopy (ATR-FTIR). ATR-FTIR captures the vibrational information of various biomolecules in plasma, thereby obtaining spectral data reflecting the overall molecular composition and structure of the sample. This method is more comprehensive than traditional biochemical testing and can simultaneously obtain information on multiple potential biomarkers.
[0037] After acquiring the raw spectral data, the module calculates an influence score for each plasma sample spectrum. This innovative metric characterizes the similarity between a sample spectrum and plasma spectra from patients with MCI, AD, non-AD, and healthy subjects. It is calculated by combining the similarity of the target sample spectrum with samples of the same type and the difference with samples of different types. This scoring mechanism effectively identifies spectral features that exhibit significant differences between diseased and healthy subjects.
[0038] Based on the calculated influence scores, the module identifies the plasma sample spectra with the highest scores as a digital marker ranking cluster. This cluster contains the spectral information that best reflects disease characteristics, laying the foundation for subsequent disease classification. The digital markers selected in this way are not only statistically significant but also have potential biological significance, as they may correspond to changes in specific biomolecules or metabolites associated with the disease.
[0039] This impact score-based screening method offers multiple advantages. First, it overcomes the limitations of traditional statistical methods (such as t-tests) that only consider differences in a single wavenumber, enabling a more comprehensive assessment of the diagnostic value of spectral features. Second, this method improves the stability and reproducibility of the selected markers because it considers both similarities within samples and differences between samples. Finally, by focusing on spectral features with high impact scores, the data dimensionality in subsequent analyses can be greatly reduced, improving computational efficiency.
[0040] Based on the above embodiment, as an optional embodiment, spectra of multiple plasma samples from patients with MCI, AD, non-AD patients, and HC are collected, including: Multiple plasma samples were collected from patients with MCI, AD, non-AD and HC, and the plasma samples were analyzed at 4000-600 cm -1 Within the wavenumber range, a preset number of cumulative spectral scans are performed at a preset resolution to obtain spectra of multiple plasma samples of MCI, AD, non-AD patients, and HC.
[0041] Among them, 4000-600cm -1 The wavenumber range refers to the mid-infrared spectral region, which is the spectral region most commonly used to study the structure and vibration modes of organic molecules. In the examples of this application, 4000-600 cm -1 The wavenumber range can be understood as the range of ATR-FTIR spectrum scanning, which is used to obtain the vibration information of various biological molecules (such as proteins, lipids, nucleic acids, etc.) in plasma samples, thereby reflecting the overall molecular composition and structure of the sample.
[0042] The preset resolution can be understood as the resolution setting used by the ATR-FTIR spectrometer when performing vibrational spectral scanning on plasma samples. It is used to determine the resolving power and data volume of the spectral data. Spectral resolution reflects the spectrometer's ability to distinguish between two adjacent spectral lines. The higher the resolution, the smaller the distance between adjacent spectral lines that can be resolved, but it also significantly increases the amount of spectral data collected. While too low a resolution may reduce the data volume, it will lose detailed information and make it impossible to accurately distinguish some key structural information in the sample.
[0043] For example, in the embodiment, 4 cm is set -1 The preset resolution is a relatively moderate resolution value, which can ensure the ability to distinguish the vibration information of the sample while avoiding the inefficiency of subsequent calculation and modeling caused by excessive data volume. -1 The resolution can well balance the fidelity of spectral information and the size of data, which is conducive to the subsequent efficient extraction of disease-related biomarker information from plasma spectral data and the establishment of identification devices covering all stages of the disease.
[0044] Cumulative spectral scanning with a preset number of scans involves performing multiple spectral scans on the same sample and averaging the spectral data from each scan. The primary purpose of cumulative scanning is to improve the signal-to-noise ratio (SNR) of the spectral data. Due to the complexity of biological samples and the influence of the measurement environment, spectral data from a single scan often contains some noise, resulting in a low SNR. However, by performing multiple scans and averaging on the same sample, random noise can be effectively eliminated, improving the reliability of the spectral data.
[0045] For example, in the embodiment of the present application, the preset number of times can be set to 32. By 32 cumulative averages, the signal of the sample itself will be retained and enhanced, while random noise will be greatly offset. The final spectral data thus obtained can more realistically reflect the inherent spectral fingerprint characteristics of the sample.
[0046] Based on the above embodiment, as an optional embodiment, the step of calculating the influence score of the spectrum of each plasma sample and determining the spectrum of the plasma sample with the highest influence score as the digital marker ranking cluster may further include the following steps: For the spectrum of any plasma sample in the spectral information space, the similarity between the spectrum of the plasma sample and the spectra of plasma samples of the same type, and the difference between the spectrum of the plasma sample and the spectra of plasma samples of different types are calculated respectively.
[0047] Specifically, for any plasma sample spectrum in the spectral information space, its similarity with the spectra of plasma samples of the same type is first calculated. The same type refers to samples belonging to the same disease category (e.g., MCI, AD, non-AD) or healthy controls (HC). Similarity can be calculated using a variety of methods, such as Euclidean distance, Pearson correlation coefficient, or cosine similarity. Its main purpose is to quantify the internal consistency between samples of the same type, which helps identify stable and representative spectral features.
[0048] At the same time, the degree of difference between the plasma sample spectrum and the spectra of plasma samples of different types must be calculated. Different types refer to samples belonging to other disease categories or healthy controls. Distinction can also be calculated using distance metrics or correlation analysis, but the goal is to quantify the degree of differentiation between samples of different categories. This helps identify spectral features that can effectively distinguish between different disease states.
[0049] By simultaneously considering both similarity and difference, the diagnostic potential of each spectral signature can be more comprehensively assessed. Ideally, a digital biomarker should exhibit high similarity within samples of the same type while exhibiting significant differences between samples of different types. This dual consideration effectively mitigates the impact of individual differences and measurement error, improving the stability and reliability of the selected digital biomarkers.
[0050] The sum of the similarity and difference of the spectra of the plasma samples is used as the influence score, and the spectra of multiple plasma samples ranked at the top in the influence score are determined to be the plasma full-spectrum digital marker ranking cluster.
[0051] Specifically, for each plasma sample spectrum, its similarity with samples of the same type and its difference with samples of different types are first calculated. Similarity reflects the consistency of spectral features within the same disease category (such as MCI, AD, non-AD) or healthy control group (HC), while difference quantifies the ability of the feature to distinguish different disease states. The combined score obtained by adding these two indicators is the influence score of the plasma sample spectrum. The above method cleverly balances the stability and distinguishing ability of the features, so that the screened markers have good repeatability and can effectively distinguish different disease states.
[0052] By adopting the above technical solution, the internal consistency and external distinguishing ability of spectral features are comprehensively considered, avoiding the bias that may be caused by a single indicator. For example, considering only similarity may select features that are highly consistent across samples of the same type but cannot distinguish diseases, while considering only difference may select features that are unstable but happen to show differences between samples. Secondly, the above method can adapt to different types of spectral data distribution characteristics. For some disease markers, they may show high similarity within the same sample; for other markers, they may mainly show significant differences between different categories. By combining these two aspects, various types of potential diagnostic markers can be captured.
[0053] Based on the calculated influence scores, the spectra of all plasma samples are ranked, and the spectra of the top-ranked samples are selected to form a ranked cluster of plasma full-spectrum digital markers. This cluster contains the spectral features with the greatest diagnostic potential, which are not only highly consistent across similar samples but also effectively distinguish between different disease states. The selection of multiple top-ranked samples rather than a single top-scoring sample ensures biomarker diversity and robustness, avoiding the risks of over-reliance on a single feature.
[0054] The plasma candidate digital marker selection module is used to iterate the plasma spectral samples in the initial disease classification device based on the plasma full-spectrum digital marker ranking cluster to obtain the plasma candidate digital marker cluster.
[0055] Building on the successful construction of a ranked cluster of plasma full-spectrum digital markers by the Plasma Sample Spectrum Acquisition Module, this example introduces a Plasma Candidate Digital Marker Selection Module, designed to further optimize and refine the digital marker set to improve the accuracy and efficiency of disease classification. This module's core task is to iteratively analyze plasma spectral samples from the initial disease classification device based on the acquired ranked cluster of plasma full-spectrum digital markers, ultimately generating a more refined and efficient cluster of plasma candidate digital markers.
[0056] Specifically, the markers in the plasma full-spectrum digital marker ranking cluster are first ranked from high to low according to their influence score. Then, starting with the marker with the highest score, they are input one by one into the initial disease classification device. The initial disease classification device can be a classification model built based on a machine learning algorithm (such as random forest or support vector machine), whose initial parameters are set based on prior knowledge or preliminary experimental results.
[0057] Each time a new marker is introduced, the performance of the classification device is evaluated. Evaluation metrics can include accuracy, sensitivity, specificity, and the area under the ROC curve (AUC). By comparing the changes in the performance of the classification device before and after the addition of a new marker, it can be determined whether the marker has made a significant contribution to improving the classification effect. If the classification performance is significantly improved after the addition of a marker, it will be retained in the candidate set; conversely, if the addition of a marker does not bring about a significant improvement, or even causes a decrease in performance, it will be removed from the candidate set.
[0058] The advantage of this iterative process is that it dynamically evaluates the contribution of each marker in the actual classification task, rather than relying solely on static statistical analysis. This approach can effectively identify markers that appear theoretically important but have limited contributions in actual classification, thereby constructing a more streamlined and efficient candidate marker cluster.
[0059] The iterative process continues until a pre-defined termination condition is met. These conditions may include: achieving a predetermined performance metric (e.g., accuracy reaching a certain threshold); reaching a pre-defined upper limit for the number of candidate markers; or failing to achieve significant performance improvement after multiple iterations. The resulting cluster of candidate plasma digital markers will contain those that perform best in the actual classification task.
[0060] By adopting the above technical solution, the practicality and reliability of digital biomarkers are significantly improved. By validating the contribution of each biomarker in real-world classification tasks, it is ensured that the final set of biomarkers is not only statistically significant but also performs well in real-world applications. Secondly, this method can significantly reduce the feature dimensionality in subsequent analyses, improving computational efficiency. By eliminating redundant or invalid biomarkers, a more streamlined but equally efficient classification model can be constructed. Furthermore, this method enhances the interpretability of the model. Each biomarker retained in the candidate set has been validated in real-world classification tasks, thereby enabling a better understanding and explanation of its role in disease diagnosis.
[0061] Finally, this iterative selection approach offers the potential for personalized medicine. By analyzing the differences in the expression of these high-efficiency markers across different patient groups, it is possible to identify disease subtypes or distinct progression patterns, thereby providing a basis for developing personalized diagnosis and treatment strategies.
[0062] The pathogenic protein digital marker cluster acquisition module is used to obtain spectral markers corresponding to the spectral absorption peaks of multiple pathogenic proteins related to MCI, AD and non-AD patients, and obtain pathogenic protein digital marker clusters.
[0063] The pathogenic protein digital marker cluster acquisition module aims to combine biological knowledge with spectral analysis technology to enhance the biological significance and diagnostic value of digital markers. The module's main task is to obtain spectral absorption peaks of multiple pathogenic proteins associated with mild cognitive impairment (MCI), Alzheimer's disease (AD), and non-AD patients, convert them into corresponding spectral markers, and ultimately construct a pathogenic protein digital marker cluster.
[0064] Furthermore, it combines known disease biological mechanisms with spectral data analysis to provide a biological basis for the screening of digital biomarkers. Traditional spectral analysis methods often focus only on statistically significant differences while ignoring the biological significance of these differences. By focusing on pathogenic proteins directly associated with the disease, it ensures that the digital biomarkers screened are not only statistically significant but also have a clear biological explanation.
[0065] Specifically, the first step is to identify key pathogenic proteins associated with MCI, AD, and non-AD patients. These proteins may include β-amyloid (Aβ), tau, phosphorylated tau (p-tau), and proteins associated with other neurodegenerative diseases. These proteins were selected based on their key roles in the development and progression of the disease and their potential as biomarkers, which have been validated in clinical studies.
[0066] Next, the purified pathogenic proteins were subjected to ATR-FTIR spectroscopy. By comparing the spectra of different proteins, the unique spectral absorption peaks of each protein can be identified. These absorption peaks usually correspond to specific structures or chemical bonds of proteins. For example, the β-folded structure of Aβ may be at 1620-1640 cm -1 The characteristic absorption peaks are generated in the range of 1000-1100 cm-1, while the phosphorylation modification of tau protein may be -1 Characteristic peaks are displayed within the range.
[0067] Then, the characteristic absorption peaks are converted into corresponding spectral markers. This step involves determining the exact wavenumber range of each absorption peak and defining it as a discrete spectral marker. For example, if Aβ -1 There is a characteristic peak at 1625-1630cm -1 The range is defined as a spectral marker.
[0068] Finally, all the spectral markers extracted from the pathogenic proteins are aggregated to form a pathogenic protein digital marker cluster. This cluster contains characteristic spectral information of multiple pathogenic proteins associated with MCI, AD, and non-AD patients.
[0069] The effects of the above method are multifaceted. First, it significantly enhances the biological significance of digital markers. Each marker directly corresponds to a specific pathogenic protein or its structural features, which enables a better understanding of the biological mechanisms behind spectral changes. Secondly, the above method improves the specificity of digital markers. Since the above markers are selected based on known pathogenic proteins, they are more likely to reflect disease-specific changes rather than general physiological fluctuations. Furthermore, the above method provides the possibility for differential diagnosis of multiple neurodegenerative diseases. By including markers of multiple disease-related proteins, a diagnostic model that can distinguish between AD, MCI and other neurodegenerative diseases can be constructed.
[0070] Furthermore, this method provides a new tool for disease progression monitoring and drug development. By tracking changes in these spectral markers directly associated with disease-causing proteins, it may be possible to detect disease progression or the effectiveness of drug treatment earlier.
[0071] Finally, the construction of a cluster of digital markers for pathogenic proteins provides important input for the subsequent plasma digital marker mining module. By combining the above-mentioned markers with clear biological significance with markers previously obtained through statistical analysis, a comprehensive diagnostic model with both statistical support and biological explanation can be constructed.
[0072] Based on the above embodiment, as an optional embodiment, multiple pathogenic proteins include Aβ42, p-tau and GFAP, and the preset wave number range includes 1800-1700 cm -1 、1750-1200cm -1 and 1200-900cm -1 .
[0073] Aβ42 refers to a 42-amino acid polypeptide fragment in the amyloid-β (Aβ) family. P-tau refers to the abnormally phosphorylated form of tau. GFAP refers to glial fibrillary acidic protein, a cytoskeletal protein in astrocytes. Aβ42, p-tau, and GFAP play crucial roles in the pathogenesis of AD.
[0074] In the examples of the present application, Aβ42, p-tau, and GFAP can be understood as three representative proteins closely related to the onset of AD. Due to its strong hydrophobicity and easier aggregation, Aβ42 is more likely to aggregate abnormally in brain tissue and form amyloid plaques than other forms of Aβ polypeptides, which is one of the most significant pathological characteristics of AD. P-tau causes conformational changes in tau protein due to abnormal hyperphosphorylation, which in turn forms neurofibrillary tangles, causes disassembly of the intracellular microtubule system, and ultimately leads to loss of neuronal function. The significant increase in GFAP levels reflects the persistent neuroinflammatory response during the AD process. The pathological changes of these three proteins reflect the pathogenesis of AD from different aspects.
[0075] Aβ42, p-tau, and GFAP are included in the embodiments of the present application to explore AD-related biomarkers and pathological mechanisms at the molecular level. First, purified Aβ42, p-tau, and GFAP protein samples are prepared separately, and their molecular vibration characteristics in the mid-infrared region are determined using ATR-FTIR spectroscopy. Proteins with different conformations and secondary structures will show specific absorption peaks on the infrared spectrum. By determining the wavenumber range corresponding to the above-mentioned characteristic absorption peak, a second spectral marker related to Aβ42, p-tau, and GFAP can be obtained. Then, by analyzing the spectral changes of AD patient plasma samples at the above-mentioned specific wavenumbers, it is expected to discover molecular signals of AD-related pathological processes such as Aβ42 aggregation, abnormal phosphorylation of tau protein, and neuroinflammation, thereby providing new biomarkers for early diagnosis and pathological monitoring of AD. Finally, the spectral marker information of different proteins such as Aβ42, p-tau, and GFAP is integrated to obtain a cluster of digital markers of pathogenic proteins.
[0076] For the above three proteins, focus on 1800-1700cm -1 、1750-1200cm-1 and 1200-900cm -1 These three wavenumber ranges are selected based on their special significance in reflecting AD-related biomolecular changes. -1 The 1750-1200 cm-1 region mainly reflects the stretching vibration of lipid C=O, which may reveal abnormal lipid metabolism related to AD. -1 The range includes the amide I and II bands of proteins, especially 1600-1700 cm -1 The amide I band and 1500-1560 cm -1 The amide II band is closely related to the secondary structure of proteins, especially the β-folded structure of Aβ fibers. -1 The region may reflect information related to membrane damage and nucleic acid oxidative damage.
[0077] In practice, ATR-FTIR spectroscopy was first performed on purified Aβ42, p-tau, and GFAP proteins, focusing on the characteristic absorption peaks within the three wavenumber ranges mentioned above. These characteristic peaks were then compared with the spectra of plasma samples from AD patients to identify the spectral markers with the greatest diagnostic value. Using this method, six AD spectral digital biomarkers were successfully identified: 1358 cm -1 、1396cm -1 、1547cm -1 、1628cm -1 、1636cm -1 and 1649cm -1 , they are all located at 1750-1200cm -1 within the area.
[0078] The biological significance of the above markers is very rich. -1 The peak at 1396 cm may reflect the vibration of nucleic acid phosphates, and its reduced intensity may be related to the damage of nucleic acids caused by increased reactive oxygen species in AD patients. -1 The peak at 1547 cm is associated with protein methyl groups and plasma phospholipids, and its reduced intensity may reflect AD-related cell membrane damage. -1 、1628cm -1 、1636cm -1 and 1649cm -1 The peaks at belong to the amide I and II regions, and the decrease in their intensity may be related to the changes in plasma Aβ42 levels and changes in the protein secondary structure.
[0079] By adopting the above technical solution, a set of spectral markers with clear biological significance is provided, making diagnostic results more interpretable. Secondly, these markers directly reflect the key pathological changes of AD, including Aβ deposition, tau protein abnormalities, and neuroinflammation, and therefore may have advantages in early diagnosis and monitoring of disease progression. Furthermore, by focusing on a specific wavenumber range, the amount of data required for analysis is greatly reduced, improving computational efficiency. Finally, the above method provides a new tool for the differential diagnosis of AD and other neurodegenerative diseases, as different diseases may exhibit different spectral characteristics within the above-mentioned specific wavenumber range.
[0080] Based on the above embodiment, as an optional embodiment, the step of obtaining spectral markers corresponding to spectral absorption peaks of multiple pathogenic proteins associated with MCI, AD, and non-AD patients may further include the following steps: The spectral absorption peak positions of multiple pathogenic protein molecular structures related to MCI, AD and non-AD patients were determined as digital markers, wherein the protein molecular structure is at least one of a chemical bond and a functional group, and the peak position of the pathogenic protein digital marker cluster is negatively correlated with the biomarker levels of MCI, AD and non-AD patients.
[0081] Building on the pathogenic protein digital marker cluster acquisition module, this example further refines the spectral marker identification method, focusing on the locations of spectral absorption peaks that characterize the molecular structures of various pathogenic proteins associated with mild cognitive impairment (MCI), Alzheimer's disease (AD), and non-AD patients. The core of this method is to directly correlate the molecular structural features of proteins, particularly specific chemical bonds and functional groups, with their characteristic absorption peaks in ATR-FTIR spectra, thereby establishing a digital marker system that is both biologically meaningful and highly specific.
[0082] Protein molecular structure is the basic structural unit that makes up a protein molecule, including but not limited to chemical bonds such as peptide bonds, hydrogen bonds, and disulfide bonds, as well as functional groups such as carboxyl groups, amino groups, and sulfhydryl groups. These chemical bonds and functional groups determine the primary, secondary, and tertiary structures of a protein, thereby influencing its functions and properties.
[0083] In the present application, it can be understood that the protein molecular structure refers specifically to the chemical bonds and functional groups that can produce characteristic absorption peaks in the ATR-FTIR spectrum. For example, the C=O stretching vibration of the peptide bond is usually at 1600-1700 cm -1 The NH bending and CN stretching vibrations are at 1500-1600 cm -1The absorption peak (i.e., amide II band) is generated within the range. The above characteristic absorption peak directly reflects the structural information of the protein and can therefore serve as an important basis for identifying and analyzing disease-related proteins.
[0084] The importance of this method lies in its ability to directly link macroscopic spectral changes with microscopic molecular structural changes. In neurodegenerative diseases such as AD, pathogenic proteins often undergo structural changes, such as misfolding or abnormal aggregation. These structural changes are directly reflected in the vibrational properties of specific chemical bonds or functional groups, resulting in characteristic absorption peaks at specific wavenumber positions in the ATR-FTIR spectrum. By determining the positions of these characteristic absorption peaks as digital markers, molecular-level changes during the disease process can be directly monitored.
[0085] Specifically, the first step is to obtain purified samples of target disease-associated proteins. These purified protein samples are scanned using ATR-FTIR spectroscopy to obtain spectra in the mid-infrared region. These spectra are then analyzed to identify characteristic absorption peaks corresponding to specific chemical bonds and functional groups in the protein. For example, the amide I band primarily reflects the C=O stretching vibration of the protein, while the amide II band primarily reflects NH bending and CN stretching vibrations. The locations of these characteristic absorption peaks are then identified as members of the pathogenic protein digital marker cluster.
[0086] It is worth noting that the examples of the present application also found that the peak position of the pathogenic protein digital marker cluster is negatively correlated with the biomarker levels of patients with different symptom stages of the target disease. This finding has important biological significance and clinical application value. The negative correlation means that as the disease progresses, the levels of certain biomarkers in the patient's body increase, while the corresponding spectral absorption peak intensity shows a downward trend. The above phenomenon may reflect the change in protein conformation or degradation process during the progression of the disease. For example, in patients with Alzheimer's disease, as the disease worsens, the levels of disease-related biomarkers such as Aβ42, p-tau181 and GFAP in plasma usually increase, while the corresponding spectral absorption peak intensity may decrease, which may be due to the aggregation or structural changes of the above proteins in plasma.
[0087] The plasma digital marker mining module is used to determine the plasma digital markers used for disease classification based on the digital logical combination of plasma candidate digital marker clusters and pathogenic protein digital marker clusters.
[0088] The plasma digital marker mining module is the core component of this device. It cleverly integrates clusters of candidate plasma digital markers screened using statistical methods with clusters of pathogenic protein digital markers with clear biological significance to determine the plasma digital markers ultimately used for disease classification. This integration method aims to improve the reliability and specificity of the markers while enhancing their biological interpretability, thereby providing a more comprehensive and accurate basis for the diagnosis of Alzheimer's disease (AD) and other neurodegenerative diseases.
[0089] Specifically, the module first performs a digital logical combination of the plasma candidate digital marker cluster and the pathogenic protein digital marker cluster. The digital logical combination can be understood as a series of set operations and logical operations. For example, the intersection of the two clusters can be found to find high-confidence markers that are identified by both methods; then the remaining markers are evaluated and those that appear in only one cluster but have important statistical or biological significance are selected. The above combination method can identify markers that are both statistically significant and have clear biological significance. These markers are likely to be key indicators for disease diagnosis. At the same time, by comprehensively considering markers from both sources, the limitations of a single method can be compensated, such as retaining certain markers that are less statistically significant but have important biological functions, thereby providing more comprehensive disease information.
[0090] Based on the digital logic combination, the module will further screen and optimize the combined markers. This includes evaluating the diagnostic efficacy of each marker (such as sensitivity, specificity, etc.), examining the correlation between markers to avoid information redundancy, and verifying the stability of markers in different patient groups. Through the above steps, a set of plasma digital markers that have high diagnostic value and can reflect the nature of the disease are finally determined. It is worth noting that the above integration method is not only applicable to the diagnosis of AD, but can also be extended to the differential diagnosis of other neurodegenerative diseases. By analyzing the relationship between different disease-specific pathogenic protein markers and statistically screened markers, it is possible to find some characteristic combinations that can effectively distinguish AD, mild cognitive impairment (MCI) and other types of dementia.
[0091] The effects brought about by the above method are multifaceted: first, it improves the accuracy and reliability of diagnosis, because the markers finally selected have both statistical support and biological basis. Secondly, it enhances the interpretability of diagnostic results, allowing clinicians to better understand the basis of diagnosis. Furthermore, the above method provides the possibility for personalized medicine. By analyzing the patient's performance on the above multidimensional markers, it is possible to identify subtypes of the disease or different progression patterns. Finally, the above integrated method also provides ideas for the discovery of new biomarkers, which may reveal some molecules or pathways that have not been noticed before but play an important role in the disease process. Through the above comprehensive and systematic plasma digital marker mining process, this device can provide strong support for the early diagnosis, disease classification and personalized treatment decisions of AD and related neurodegenerative diseases.
[0092] The disease classification device construction module is used to train the iterated initial disease classification device based on the plasma digital markers to obtain a trained disease classification device.
[0093] Based on the above embodiment, as an optional embodiment, a disease classification model device evaluation module is used to evaluate the performance of the disease classification device based on the output results of the disease classification device.
[0094] Based on the above embodiment, as an optional embodiment, the process of training the iterative initial disease classification device according to the plasma digital markers to obtain the trained disease classification device may further include the following steps: The plasma digital markers are divided into training set, validation set and test set. The iterative initial disease classification device is trained through the training set, the parameters of the trained initial disease classification device are adjusted through the validation set, and the performance of the trained initial disease classification device is evaluated through the test set and the disease classification model device evaluation module to obtain a trained disease classification device.
[0095] Each spectral feature is divided into training set, validation set and test set, and the initial plasma digital marker mining and disease classification device of the target spectral marker cluster is constructed by random forest algorithm.
[0096] Specifically, during the machine learning modeling process, a training set is required for parameter estimation and device fitting, while a validation set and a test set are also needed to evaluate the device's generalization and predictive performance to avoid overfitting. Therefore, all samples must be divided into three subsets: training, validation, and test sets, according to a specific ratio. A common practice is to use the Kennard-Stone or other specialized sample partitioning algorithms, distributing samples in a ratio of 7:1.5:1.5 or similar, ensuring that the three subsets cover the sample distribution while avoiding overfitting.
[0097] The initial plasma digital marker mining and disease classification device is trained through the training set, and the parameters of the trained initial plasma digital marker mining and disease classification device are evaluated and adjusted through the validation set to obtain the trained plasma digital marker mining and disease classification device.
[0098] Specifically, the training set is fed into the initial plasma digital marker mining and disease classification device, which is then fully trained using a random forest algorithm. During the training process, the algorithm automatically adjusts the internal parameters of the initial plasma digital marker mining and disease classification device based on the training data, achieving optimal classification performance on the training set.
[0099] Next, the trained initial plasma digital marker mining and disease classification device can be applied to a validation set to evaluate its generalization and predictive performance on new data. Further evaluation metrics for the device on the validation set need to be calculated. This analysis provides a comprehensive understanding of the device's performance in identifying patients at different symptom stages and healthy individuals. If the device performs poorly in certain areas, such as a low detection rate for early-stage patients or a high misclassification rate for healthy individuals, further adjustment and optimization of the device parameters are necessary.
[0100] Based on the validation set evaluation, the plasma digital biomarker mining and disease classification device construction module analyzes various factors that affect device performance, such as algorithm hyperparameter settings, feature selection strategies, and the quality of training data, and optimizes these factors to achieve optimal device performance. For example, parameters such as the number of trees and maximum tree depth in the random forest algorithm can be adjusted; a more influential subset of spectral features can be reselected as input; or preprocessing such as feature expansion and data cleaning can be performed on the training data.
[0101] After one or more cycles of training-validation-parameter adjustment, when the evaluation indicators of the device on the validation set meet the expected requirements, the final parameters of the device can be determined, and a trained plasma digital marker mining and disease classification device can be obtained.
[0102] The performance of the plasma digital marker mining and disease classification device is evaluated using a test set. Specifically, the following steps may be involved: The spectral features of the samples in the test set were input into the plasma digital biomarker mining and disease classification system to obtain recognition results. Based on the recognition results, the sensitivity and specificity of the plasma digital biomarker mining and disease classification system were calculated. A receiver operating characteristic (ROC) curve was generated based on the sensitivity and specificity, and the area under the ROC curve (AUC) was calculated. The performance of the plasma digital biomarker mining and disease classification system was evaluated based on sensitivity, specificity, and AUC.
[0103] Specifically, the retained test set can be fed into the trained plasma digital marker mining and disease classification device for testing. For each sample in the test set, the device outputs an identification result based on its spectral feature vector, determining whether the sample is diseased or healthy.
[0104] Next, two key evaluation metrics need to be calculated based on the device's recognition output: sensitivity and specificity. Sensitivity refers to the proportion of positive samples (disease samples) correctly identified by the device, reflecting its ability to detect patients; specificity refers to the proportion of negative samples (healthy samples) correctly identified by the device, reflecting its ability to exclude healthy individuals. The specific process for calculating sensitivity and specificity is to compare the device's predictions with the known true categories of the test set samples one by one, and then count the number of true positive (TP), false positive (FP), true negative (TN), and false negative (FN) samples. Sensitivity is then calculated using the formulas: TP = TP / (TP + FN) and specificity = TN / (TN + FP).
[0105] The sensitivity and specificity of the device on the test set are used as input to generate the receiver operating characteristic (ROC) curve, and the area under the ROC curve (AUC) is calculated. The ROC curve is a visualization method that reflects the classification performance of a binary classifier. The horizontal axis is the false positive rate (1-specificity), and the vertical axis is the true positive rate (sensitivity). The ideal classifier should make the ROC curve as close to the upper left corner as possible. The closer the corresponding AUC value is to 1, the better the performance of the classifier. Calculating the AUC can not only quantify the overall performance of the classifier, but also objectively evaluate its balance between detecting diseased samples and excluding healthy samples.
[0106] Finally, the diagnostic performance of the plasma digital biomarker mining and disease classification device was comprehensively evaluated and analyzed based on a series of evaluation metrics, including sensitivity, specificity, and AUC. For example, if the device has high sensitivity but low specificity, it indicates that it performs well in detecting patients, but may also have a high rate of false positives in healthy individuals, requiring appropriate adjustments. An AUC value close to 1 indicates that the device has excellent classification capabilities, effectively identifying diseased samples at different symptom stages while also effectively excluding healthy samples. By analyzing these metrics, we can comprehensively evaluate the device's potential for various applications, including early disease detection, clinical diagnosis, and disease progression monitoring.
[0107] Based on the above embodiment, as an optional embodiment, the process of iterating the plasma spectrum samples in the initial disease classification device based on the plasma full spectrum digital marker ranking cluster to obtain the plasma candidate digital marker cluster may further include the following steps: According to the influence score of the digital markers in the plasma full-spectrum digital marker ranking cluster, the digital markers are input into the initial disease classification device in sequence, and the impact of each digital marker on the initial disease classification device is evaluated through the disease classification model device evaluation module to obtain the evaluation results.
[0108] Based on the evaluation results, no less than 20% of the digital markers in the plasma candidate digital marker cluster are screened out.
[0109] During the process of mining plasma digital markers and building a disease classification device, this example introduces a disease classification model device evaluation module to further optimize the cluster of candidate plasma digital markers and improve the performance of the disease classification device. This improvement aims to identify the most diagnostically valuable digital markers through dynamic evaluation and iterative optimization, thereby improving the accuracy and reliability of diagnosing Alzheimer's disease (AD) and other neurodegenerative diseases.
[0110] Specifically, the plasma spectral samples in the initial disease classification device are iterated based on the ranked clusters of plasma full-spectrum digital markers. This process begins by sequentially inputting digital markers into the initial disease classification device according to their influence scores within the ranked clusters of plasma full-spectrum digital markers. The influence scores reflect the potential importance of each digital marker in distinguishing different disease states, so inputting them in this order prioritizes markers that are likely to contribute significantly to the classification effect.
[0111] After each new digital marker is input, the disease classification model device evaluation module assesses the marker's impact on the performance of the initial disease classification device, generating an evaluation result. This evaluation process may include calculating changes in metrics such as classification accuracy, sensitivity, specificity, and AUC, as well as examining whether the marker has resulted in a statistically significant performance improvement. The advantage of this dynamic evaluation method is that it can verify the contribution of each marker in the actual classification task, rather than relying solely on static statistical analysis.
[0112] Based on the above evaluation results, this example further eliminated no less than 20% of the digital markers in the cluster of candidate plasma digital markers. This ratio aims to strike a balance between retaining sufficient information and improving model efficiency. Eliminating 20% or more of the digital markers significantly reduces feature dimensionality, thereby reducing computational complexity and the risk of overfitting. At the same time, 80% or less of the high-quality markers are still retained, ensuring that the model's diagnostic performance is not significantly reduced due to information loss.
[0113] Second, this ratio takes into account the complexity and redundancy of biological systems. In complex biological systems, multiple markers may reflect similar biological processes or provide overlapping information. By eliminating at least 20% of the markers, this redundant information can be effectively removed while retaining sufficient diversity to capture different aspects of the disease.
[0114] It's important to note that while a minimum 20% screening ratio is set, the exact percentage to be screened out is determined dynamically based on the evaluation results of the disease classification model device evaluation module. For example, if the evaluation finds that 25% of the markers do not contribute significantly to model performance, these markers should be screened out. This dynamic performance-based adjustment ensures the scientific and effective nature of the screening process.
[0115] This iterative optimization and dynamic evaluation approach has multiple benefits. First, it significantly improves the practicality and reliability of digital biomarkers. By validating the contribution of each biomarker in real-world classification tasks, it ensures that the final set of biomarkers is not only statistically significant but also performs well in real-world applications. Second, this approach can significantly reduce the feature dimensions in subsequent analyses, improving computational efficiency. By eliminating redundant or invalid biomarkers, a more streamlined but equally efficient classification model can be constructed. Furthermore, this approach enhances the interpretability of the model. Each biomarker retained in the candidate set has been validated in real-world classification tasks, allowing for a better understanding and explanation of its role in disease diagnosis.
[0116] Based on the above embodiment, as an optional embodiment, the performance of the disease classification device includes at least one of specificity, sensitivity, and the area of the ROC curve.
[0117] The specificity and sensitivity of the initial disease classification device after iteration are not less than 80%, and the area under the ROC curve is greater than 75%.
[0118] Among them, specificity and sensitivity are two key indicators for evaluating the performance of diagnostic tests. Specificity reflects the ability of the disease classification device to correctly identify healthy individuals, while sensitivity reflects its ability to correctly identify diseased individuals. In this embodiment, the specificity and sensitivity of the initial disease classification device after iteration are set to be no less than 80%. The setting of this standard is based on an in-depth understanding of clinical practice needs. The 80% threshold represents a relatively high diagnostic accuracy, which can greatly reduce the risk of misdiagnosis and missed diagnosis. A specificity of 80% means that among healthy people, 80% can be correctly identified as healthy, greatly reducing the possibility of overdiagnosis. Similarly, a sensitivity of 80% ensures that most individuals with AD or other related neurodegenerative diseases can be discovered in a timely manner, providing the possibility for early intervention.
[0119] In addition to specificity and sensitivity, this embodiment also introduces the ROC curve and its area (AUC) as evaluation indicators. The ROC curve is an important tool to reflect the performance of the classification model. It intuitively shows the performance of the model under different decision thresholds. The AUC value ranges from 0 to 1, where 0.5 represents random guessing and 1 represents perfect classification. In this embodiment, the area of the ROC curve is required to be greater than 75%. The setting of this standard is also based on high requirements for model performance. An AUC greater than 75% means that the model has good discrimination ability and can maintain a high diagnostic accuracy under various decision thresholds. This is particularly important for actual clinical applications because it ensures the stable performance of the model under different sensitivity requirements.
[0120] During model optimization, special attention was paid to balancing specificity and sensitivity. This is because in actual clinical applications, these two metrics often present a trade-off. Improving specificity may reduce sensitivity, and vice versa. By fine-tuning model parameters and decision thresholds, we strive to find the optimal balance between these two metrics, meeting the dual requirement of no less than 80%. At the same time, by adjusting model complexity and adding regularization, we strive to increase the area under the receiver operating characteristic (ROC) curve to achieve a standard of greater than 75%.
[0121] By employing these technical solutions, the disease classification device ensures high reliability in real-world clinical applications. Specificity and sensitivity exceeding 80% indicate that the device provides reliable results, both in excluding healthy individuals and identifying diseased individuals. Furthermore, an AUC greater than 75% demonstrates that the device maintains stable performance across different decision thresholds, providing clinicians with greater flexibility to tailor diagnostic strategies to specific circumstances.
[0122] Based on the above embodiment, as an optional embodiment, the specificity and sensitivity of the initial disease classification device after training are not less than 85%, and the area of the ROC curve approaches 1.
[0123] During the development of the plasma digital marker mining and disease classification device, this example set a higher performance target to further improve the diagnostic performance of the disease classification device: the specificity and sensitivity of the initial disease classification device after training should be no less than 85%, and the area under the receiver operating characteristic (ROC) curve should approach 1. This target aims to build a high-performance model with excellent performance in the diagnosis of Alzheimer's disease (AD) and other neurodegenerative diseases, meeting the high diagnostic accuracy requirements in clinical practice.
[0124] Specifically, early diagnosis of AD is crucial for timely intervention and treatment, while misdiagnosis or missed diagnosis can have a serious impact on patients' quality of life and treatment outcomes. A specificity of over 85% means the model can more accurately identify healthy individuals, significantly reducing false-positive results and avoiding unnecessary anxiety and further examinations. Similarly, a sensitivity of over 85% ensures that the vast majority of AD patients can be detected promptly, creating conditions for early intervention. An ROC curve area approaching 1 indicates that the model maintains extremely high classification performance at various decision thresholds, which is crucial for adapting to different clinical scenarios and personalized diagnostic strategies.
[0125] Based on the above examples, the following will describe the training process of the plasma digital marker mining and disease classification device in combination with actual application scenarios. Figure 2 , Figure 2 A flowchart of the plasma digital marker mining and disease classification device training process provided in an embodiment of the present application.
[0126] First, we selected three indicative biomarkers, GFAP, p-tau, and Aβ42, which are closely related to the pathogenesis of AD, and measured their ATR-FTIR spectra. Based on the characteristic absorption peaks of the spectra of the above indicative biomarkers, we determined 20 core spectral digital biomarkers, including 1329 cm -1 、1358cm -1 、1396cm -1 、1409cm -1 、1452cm -1 、1456cm -1 、1503cm -1 、1547cm -1 、1622cm -1 、1628cm -1 、1636cm -1 、1649cm -1 、2852cm -1 、2885cm -1 、2926cm -1、2936cm -1 、3182cm -1 、3260cm -1 、3274cm -1 and 3288cm -1 .
[0127] Then, plasma samples covering patients at different symptom stages of AD and age-matched healthy controls were collected. Specifically, they included: Cohort 1: 275 AD patients and 189 HC from the clinic; Cohort 2: 18 AD patients and 344 HC from the community; Cohort 3: 151 MCI patients; Cohort 4: 106 Dementia with Lewy Bodies (DLB) patients, 106 Frontotemporal Dementia (FTD) patients and 135 Progressive Supranuclear Palsy (PSP) patients. The characteristics of the recruited patients are shown in Table 1. Table 1. Statistics of enrolled patients' characteristics All plasma samples were spectrally scanned and preprocessed using ATR-FTIR technology to obtain enhanced spectral data, such as second-order derivative spectra. The data were randomly divided into three subsets: training, validation, and test sets in a ratio of 7:1.5:1.5.
[0128] Next, the Relief-F algorithm was used to calculate the influence score of each spectral feature on the training set and validation set, and the 27 with the highest scores were selected as the spectral features based on machine learning. These 27 digital biomarkers based on machine learning were logically ANDed with the previous 20 core spectral digital biomarkers as follows: Figure 3 As shown, 6 AD spectrum digital biomarkers were obtained, namely 1358cm -1 、1396cm -1 、1547cm -1 、1628cm -1 、1636cm -1 and 1649cm -1 Its biochemical significance is shown in Table 2. Table 2. Plasma digital markers and their biochemical significance Please refer to Figure 4 , Figure 4ATR-FTIR spectrum comparison and diagnostic results diagram of AD and HC provided in the embodiment of this application. -1 、1628cm -1 and 1636cm -1 The peaks were selected based on the spectrum of purified Aβ42; 1396 cm -1 、1636cm -1 and 1649cm -1 The peak selected was based on the spectrum of purified p-tau; 1358 cm -1 and 1649cm -1 was selected based on the spectral peak of purified GFAP (e.g. Figure 4 (as shown in A).
[0129] These six AD spectral digital biomarkers were input into the random forest algorithm, the initial AD recognition device was trained using the training set, and the device parameters were optimized and adjusted using the validation set to obtain a trained AD recognition device.
[0130] The performance of the AD recognition device was evaluated on the reserved test set. First, the average spectra of AD patients and HC were compared, and it was found that the absorbance of the average spectrum of AD patients was lower than that of HC (e.g. Figure 4 The above six AD spectral digital biomarkers were significantly decreased in the AD group (P < 0.05, as shown in Figure 2). Figure 4 B).
[0131] The evaluation index of the complex disease screening and auxiliary diagnosis model based on plasma spectra constructed in this application can be defined by the confusion matrix. The confusion matrix H is expressed as: Among them, TP: is judged as positive, but is actually positive, that is, true positive; FN: is judged as negative, but is actually positive, that is, false negative; FP: is judged as positive, but is actually negative, that is, false positive; TN: is judged as negative, but is actually negative, that is, true negative; Indicators in the field of disease detection include sensitivity and specificity. Sensitivity refers to the probability of not missing a diagnosis during screening and diagnosis, while specificity refers to the probability of not misdiagnosing a disease during screening and diagnosis, expressed as: In the evaluation of the Alzheimer's disease screening and diagnosis model, two evaluation indicators, sensitivity and specificity, are used to objectively evaluate the application performance of the model. The results of this example are as follows: Table 3. Results of intelligent detection of Alzheimer's disease Alzheimer's disease symptom stages Sensitivity % Specificity% AD vs HC 88.8% 86.4% MCI vs HC 80.95% 85.29% DLB vs HC 88.5% 75% FTD vs HC 80.8% 77.1% PSP vs HC 75.8% 72.6% AD vs MCI 77.93% 80.95% AD vs DLB 77.2% 74.6% AD vs FTD 74.2% 72.4% AD vs PSP 76.1% 75.7% The results showed that the mined plasma digital markers not only have the ability to screen and assist in diagnosis of the MCI stage, but can also perform preliminary differentiation between AD and non-AD dementia (DLB, FTD and PSP).
[0132] The diagnostic effects of these six AD spectral digital biomarkers on AD were further evaluated on the test set (e.g. Figure 4 C). The results showed that the diagnostic efficiency of a single absorption peak was poor, but after combining the six biomarkers, the sensitivity for diagnosing AD was 88.2%, the specificity was 84.1%, and the AUC for distinguishing AD from HC reached 0.92 (as shown in Figure 4 C).
[0133] For further validation, the device was tested using an independent cohort 2. The results showed that the device had a sensitivity of 88.8%, a specificity of 86.4%, and an AUC of 0.94 in differentiating AD from HC ( Figure 4 D), further confirming the excellent ability of this ATR-FTIR spectroscopy device in diagnosing AD.
[0134] The specificity figures of 84.1% and 86.4% reflect the ATR-FTIR spectroscopy device's ability to correctly identify healthy controls (HC) across different test sets, demonstrating the device's performance stability and generalizability. The 84.1% specificity on the initial test set indicates the device's ability to accurately exclude a majority of healthy samples. In an independent validation cohort 2, the specificity slightly increased to 86.4%, further confirming the device's effectiveness. This stable and slightly improving performance across different datasets demonstrates not only the device's good reproducibility but also its high diagnostic performance with new, independent samples. This slight increase in specificity may be due to the device's good generalization or differences in sample characteristics within cohort 2. Regardless, the high specificity of over 80% achieved on both test sets highlights the device's potential in clinical applications to reduce misdiagnosis and unnecessary treatment. These results strongly support the promising application of ATR-FTIR spectroscopy in AD diagnosis and demonstrate its potential as a reliable diagnostic tool.
[0135] The device was also evaluated for its ability to differentiate between MCI and HC, and AD. The results showed that the device had an AUC of 0.89 for differentiating between MCI and HC, and an AUC of 0.86 for differentiating between MCI and AD.
[0136] Based on the above embodiment, as an optional embodiment, in addition to evaluating the performance of the device in diagnosing AD, this embodiment further evaluates the ability of ATR-FTIR spectroscopy technology in detecting MCI and distinguishing MCI from HC and AD.
[0137] Please refer to Figure 5 , Figure 5 A schematic diagram of the MCI identification capability of Fourier transform infrared spectroscopy provided in the embodiment of the present application. First, ATR-FTIR spectroscopy was performed on the plasma samples of 151 MCI patients in cohort 3 and AD patients and HC in cohort 1 to obtain the corresponding second-order derivative spectra (such as Figure 5 By analyzing the above spectral data, 11 spectral digital biomarkers were selected to distinguish between MCI and HC, and 9 spectral digital biomarkers were selected to distinguish between MCI and AD. Figure 4 The second-order derivative spectra of A are marked with different colors.
[0138] Next, the performance of ATR-FTIR spectroscopy in differentiating MCI from HC and AD was evaluated using the spectral digital biomarkers screened above. Figure 5 (B) ATR-FTIR spectroscopy demonstrated a sensitivity of 88.8%, a specificity of 86.4%, and an AUC of 0.89 for differentiating between MCI and HC; and a sensitivity of 75.8%, a specificity of 72.6%, and an AUC of 0.86 for differentiating between MCI and AD. This demonstrates that ATR-FTIR spectroscopy is not only effective in diagnosing AD but also highly effective in identifying the early stages of AD, namely MCI.
[0139] In addition, the relationship between 20 spectral digital biomarkers used to diagnose MCI and the Mini-Mental State Examination scores of AD patients was further analyzed. Figure 5 C), of which 6 spectral digital biomarkers (including 1365cm -1 、1366cm -1 、1382cm -1 、1396cm -1 、1397cm -1 and 1398cm -1 ) was positively correlated with the MMSE score. The MMSE score is commonly used to assess cognitive function in AD patients, with lower scores indicating more severe cognitive impairment. This result suggests that these spectral digital biomarkers can, to a certain extent, reflect the degree of AD progression.
[0140] In summary, this example demonstrates that ATR-FTIR spectroscopy is not only effective in diagnosing AD but also highly effective in identifying early MCI stages of AD, providing new insights into early disease screening. Furthermore, some of the spectral digital biomarkers revealed by this technology have the potential to assess the progression of AD, potentially providing more comprehensive diagnostic and monitoring tools for clinical practice.
[0141] Based on the above embodiment, as an optional embodiment, this embodiment also evaluates the potential of ATR-FTIR spectroscopy to distinguish AD from other neurodegenerative diseases.
[0142] Specifically, plasma samples from 106 DLB patients, 106 FTD patients, and 135 PSP patients from cohort 4 were combined with AD patient samples from cohort 1, and their second-derivative ATR-FTIR spectra were analyzed.
[0143] Please refer to Figure 6 , Figure 6 A schematic diagram of the differential diagnosis performance of Fourier transform infrared spectroscopy provided in the embodiment of the present application. First, from the second derivative spectra of DLB patients and AD patients (such as Figure 6 A), 12 spectral digital biomarkers were screened out to distinguish DLB from AD. The results showed (as shown in Figure 6 Using these 12 biomarkers, ATR-FTIR spectroscopy achieved a sensitivity of 77.2%, a specificity of 74.6%, and an AUC of 0.83 in differentiating DLB from AD.
[0144] Secondly, from the second derivative spectra of FTD patients and AD patients (such as Figure 6 C), 12 spectral digital biomarkers were screened out to distinguish FTD from AD. The results showed (as shown in Figure 6 Using these 12 biomarkers, ATR-FTIR spectroscopy had a sensitivity of 74.2%, a specificity of 72.4%, and an AUC of 0.80 in distinguishing FTD from AD.
[0145] Finally, from the second derivative spectra of PSP patients and AD patients (e.g. Figure 6 E), 13 spectral digital biomarkers were screened out to distinguish PSP from AD. The results showed (as shown in Figure 6 F), using these 13 biomarkers, ATR-FTIR spectroscopy had a sensitivity of 76.1%, a specificity of 75.7%, and an AUC of 0.81 in distinguishing PSP from AD.
[0146] These results demonstrate that ATR-FTIR spectroscopy is not only highly effective in diagnosing AD but also can distinguish it from other common neurodegenerative diseases, such as DLB, FTD, and PSP, demonstrating its potential for differential diagnosis of neurodegenerative diseases. This technology provides a new auxiliary tool for the accurate diagnosis of these diseases in clinical practice.
[0147] Based on the above embodiment, as an optional embodiment, in order to further demonstrate the importance of ATR-FTIR spectroscopy technology in detecting AD pathological changes, this embodiment measured the biomarker levels in plasma samples of AD patients and MCI patients in cohorts 1 and 3, and analyzed the correlation between spectral digital biomarkers and the above biomarkers.
[0148] Specifically, the levels of three biomarkers closely related to the pathogenesis of AD, p-tau181, Aβ42, and GFAP, were measured in the participants' plasma. At the same time, based on the previous ATR-FTIR spectral analysis results of purified p-tau, GFAP, and Aβ42 proteins, the 1396 cm -1 、1636cm -1 、1649cm -1 、1358cm -1 、1547cm -1 and 1628cm -1 6 spectral digital biomarkers.
[0149] Please refer to Figure 7 , Figure 7 A schematic diagram of the linear relationship between the plasma GFAP, p-Tau and Aβ42 levels of AD patients in cohort 1 and cohort 3 and the absorbance of the corresponding AD spectral digital biomarkers is provided in the embodiment of this application. Further correlation analysis found (such as Figure 7 A), 1396cm -1 The spectral digital biomarker at was significantly negatively correlated with plasma p-tau181 levels (p<0.01, r=-0.11). The absorption peak of this biomarker was selected based on the spectral peak of purified p-tau protein.
[0150] Similarly, 1358cm -1 and 1649cm -1 The spectral digital biomarker at 1358 cm was determined based on the spectral peak of purified GFAP protein. -1 There was a significant negative correlation between the plasma GFAP level (p = 0.03, r = -0.09) (e.g. Figure 7 B).
[0151] And 1547cm -1 、1628cm -1 and 1636cm -1 The spectral digital biomarkers at 4 and 8 were derived from the spectral peaks of purified Aβ42 protein, but they had no significant correlation with plasma Aβ42 levels (e.g. Figure 7 C).
[0152] These results demonstrate that some digital biomarkers in ATR-FTIR spectroscopy can indeed reflect the changing trends of key proteins (such as p-tau and GFAP) in the pathological progression of AD, providing a molecular-level biological explanation for the application of ATR-FTIR spectroscopy in AD diagnosis. Correlation analysis with protein levels not only validates the value of ATR-FTIR spectroscopy in detecting AD pathological changes but also lays the foundation for further exploration of the intrinsic connection between spectral changes and AD pathogenesis.
[0153] The present application also discloses a disease screening device, comprising a disease identification model constructed by the plasma digital marker mining and disease classification device described in any one of the above items.
[0154] The present application also discloses an electronic device comprising a memory and a processor, wherein the memory stores a disease recognition model that can be loaded and executed by the processor to implement the plasma digital marker mining and disease classification device described in any of the above items.
[0155] Among them, the electronic device can be an electronic device such as a desktop computer, a laptop computer or a cloud server, and the electronic device includes but is not limited to a processor and a memory. For example, the electronic device can also include input and output devices, network access devices and buses, etc.
[0156] The processor in the present application may include one or more processing cores. The processor calls the data stored in the memory by running or executing the instructions, programs, code sets or instruction sets stored in the memory, performs the various functions of the present application and processes data. The processor can be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller and a microprocessor. It is understandable that for different devices, the electronic device for realizing the above-mentioned processor function can also be other, and the embodiments of the present application are not specifically limited.
[0157] Among them, the memory can be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device, or it can be an external storage device of the electronic device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD) or flash memory card (FC) equipped on the electronic device. In addition, the memory can also be a combination of an internal storage unit and an external storage device of the electronic device. The memory is used to store computer programs and other programs and data required by the electronic device. The memory can also be used to temporarily store data that has been output or is to be output. This application does not impose any restrictions on this.
[0158] The present application also discloses a computer-readable storage medium that stores a disease recognition model constructed by the plasma digital marker mining and disease classification apparatus as described in any one of the above items and that can be loaded and executed by a processor.
[0159] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0160] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0161] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0162] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard drives, magnetic disks or optical disks.
[0163] The foregoing is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of the present disclosure. In other words, any equivalent variations and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the disclosure and the practical implications thereof.
[0164] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as illustrative only, and the scope and spirit of the present disclosure are to be defined by the claims.
Claims
1. A plasma digital marker mining and disease classification device, characterized in that: include: a plasma sample spectrum acquisition module, configured to collect spectra of multiple plasma samples from patients with MCI, AD, non-AD, and HC, and calculate an influence score for the spectrum of each plasma sample, and determine the spectrum of the plasma sample with the highest influence score as the digital marker ranking cluster, wherein the influence score represents the similarity of the plasma sample spectrum to the spectra of the plasma samples of MCI, AD, non-AD, and HC; a plasma candidate digital marker selection module, configured to iterate the plasma spectrum samples in the initial disease classification device based on the plasma full spectrum digital marker ranking cluster to obtain a plasma candidate digital marker cluster; A pathogenic protein digital marker cluster acquisition module is used to acquire spectral markers corresponding to spectral absorption peaks of multiple pathogenic proteins related to the MCI, AD and non-AD patients to obtain a pathogenic protein digital marker cluster; A plasma digital marker mining module, configured to determine plasma digital markers for disease classification based on a digital logical combination of the plasma candidate digital marker cluster and the pathogenic protein digital marker cluster; The disease classification device construction module is used to train the iterated initial disease classification device based on the plasma digital markers to obtain a trained disease classification device.
2. The plasma digital marker mining and disease classification device according to claim 1, characterized in that: The method includes collecting spectra of multiple plasma samples from patients with MCI, AD, non-AD patients, and HC, including: Multiple plasma samples were collected from patients with MCI, AD, non-AD and HC, and the plasma samples were analyzed at 4000-600 cm -1 Within the wavenumber range, cumulative spectral scans are performed a preset number of times at a preset resolution to obtain spectra of the multiple plasma samples of the MCI, AD, non-AD patients, and HC.
3. The plasma digital marker mining and disease classification device according to claim 1, characterized in that: Calculating the influence score of the spectrum of each plasma sample and determining the spectrum of the plasma sample with the highest influence score as the digital marker ranking cluster includes: For a spectrum of any plasma sample in the spectral information space, respectively calculating the similarity between the spectrum of the plasma sample and spectra of plasma samples of the same type, and the difference between the spectrum of the plasma sample and spectra of plasma samples of different types; The sum of the similarity and difference of the spectra of the plasma samples is used as an influence score, and the spectra of the multiple plasma samples ranked higher in influence scores are determined as a plasma full-spectrum digital marker ranking cluster.
4. The plasma digital marker mining and disease classification device according to claim 1, characterized in that: The obtaining of spectral markers corresponding to spectral absorption peaks of multiple pathogenic proteins associated with the MCI, AD, and non-AD patients includes: The spectral absorption peak positions of multiple pathogenic protein molecular structures related to the MCI, AD and non-AD patients are determined as digital markers, wherein the protein molecular structure is at least one of a chemical bond and a functional group, and the peak position of the pathogenic protein digital marker cluster is negatively correlated with the biomarker levels of the MCI, AD and non-AD patients.
5. The plasma digital marker mining and disease classification device according to claim 1, characterized in that: The multiple pathogenic proteins include Aβ42, p-tau and GFAP, and the preset wavenumber range includes 1800-1700 cm -1 、1750-1200cm -1 and 1200-900cm -1 .
6. The plasma digital marker mining and disease classification device according to claim 1, characterized in that: The plasma digital marker mining and disease classification device further comprises: a disease classification model device evaluation module for evaluating the performance of the disease classification device based on the output results of the disease classification device; The method of iterating the plasma spectrum samples in the initial disease classification device based on the plasma full spectrum digital marker ranking cluster to obtain the plasma candidate digital marker cluster includes: According to the influence scores of the digital markers in the plasma full spectrum digital marker ranking cluster, the digital markers are sequentially input into the initial disease classification device, and the impact of each digital marker on the initial disease classification device is evaluated by the disease classification model device evaluation module to obtain an evaluation result; Screening out no less than 20% of the digital markers in the plasma candidate digital marker cluster based on the evaluation results; The iterative initial disease classification device is trained based on the plasma digital markers to obtain a trained disease classification device, including: The plasma digital markers are divided into a training set, a validation set, and a test set. The iterated initial disease classification device is trained using the training set. The parameters of the trained initial disease classification device are adjusted using the validation set. The performance of the trained initial disease classification device is evaluated using the test set and the disease classification model device evaluation module to obtain a trained disease classification device.
7. The plasma digital marker mining and disease classification device according to claim 6, characterized in that: The performance of the disease classification device includes at least one of specificity, sensitivity, and the area of the ROC curve; The specificity and sensitivity of the initial disease classification device after iteration are not less than 80%, and the area under the ROC curve is greater than 75%; The specificity and sensitivity of the trained initial disease classification device are not less than 85%, and the area of the ROC curve is close to 1.
8. A disease screening device, characterized in that: Including a disease classification model constructed by the plasma digital marker mining and disease classification device as described in any one of claims 1-7.
9. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the disease classification model constructed by the plasma digital marker mining and disease classification device as described in any one of claims 1-7.
10. A computer storage medium, characterized in that The computer storage medium stores instructions, and when the instructions are executed, the disease classification model constructed by the plasma digital marker mining and disease classification device according to any one of claims 1 to 7 is executed.
Citation Information
Cited By
Disease-specific quantitative trait site recognition method based on multi-omics integration
CN122067599A
Disease-Specific Quantitative Trait Locus Identification Method Based on Multi-omics Integration
CN122067599B