Extracellular matrix aging assessment method and system based on plasma proteomics
Patent Information
- Application Number
- CN202611194843.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-22
AI Technical Summary
但其数据通常存在维度高、蛋白间相关性强、特征冗余及缺失值等问题,若缺乏有效的数据质控、缺失值处理、标准化及特征筛选机制,容易导致模型稳定性不足、预测结果偏差较大以及关键衰老标志物难以识别
[0020]本申请实施例至少包括以下有益效果:本申请提供一种基于血浆蛋白质组学的细胞外基质衰老评估方法及系统,该方案通过获取并预处理人群样本数据及血浆蛋白质组学数据,能够提高数据完整性、一致性和可比性;通过基于预设细胞外基质蛋白列表提取候选蛋白数据,使模型构建过程聚焦于与细胞外基质衰老相关的蛋白特征,降低无关蛋白干扰;通过对候选蛋白数据进行质量控制、缺失值填补及表达量标准化处理,能够减少异常值及缺失数据对评估结果的影响;通过构建细胞外基质衰老评估模型,能够从血浆蛋白表达水平中提取反映细胞外基质衰老状态的特征信息,得到细胞外基质预测年龄,实现对细胞外基质衰老程度的量化表征,并通过模型性能验证其预测准确性与泛化能力,确保评估结果具有可量化的置信度;通过对细胞外基质预测年龄与实际年龄进行偏差校正,能够降低年龄相关系统偏差对评估结果的影响,进一步得到更能反映个体衰老快慢的细胞外基质衰老年龄差;同时,结合细胞外基质预测年龄、细胞外基质衰老年龄差及模型性能指标生成综合评估结果,使评估结果兼具个体化、可量化及可验证性,从而提高细胞外基质衰老评估的准确性、稳定性和应用价值。
Smart Images

Figure CN122800249A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of aging assessment technology, and in particular to a method and system for assessing extracellular matrix aging based on plasma proteomics. Background Technology
[0002] Aging status assessment has become an important technical tool in health management, disease risk stratification, and early intervention. These assessment methods typically use biomarkers, clinical indicators, or organ function test data to construct biological age assessment models to measure the difference between an individual's actual physiological state and their chronological age. Compared to chronological age calculated solely based on birth time, biological age more accurately reflects the degree of functional decline and the risk of developing chronic diseases.
[0003] The extracellular matrix, as a crucial foundation for maintaining tissue structure, mechanical properties, and cellular microenvironment homeostasis, undergoes degenerative changes during aging, including elastic fiber breakage, abnormal collagen cross-linking, and disordered matrix metalloproteinase activity. These changes are closely associated with age-related diseases such as chronic obstructive pulmonary disease (COPD), arteriosclerosis, and osteoarthritis, making it a key structural basis for disease development and listed as the 14th major marker of aging in humans. However, current aging assessment methods primarily focus on evaluating the overall state of the body or organ function, failing to directly reflect the degree of tissue aging reflected by the structural and functional degeneration of the extracellular matrix.
[0004] Furthermore, while plasma proteomics offers advantages such as high throughput, quantifiability, and the ability to reflect the body's pathophysiological state, providing a potential technical pathway for the aging characterization of extracellular matrix-related proteins, its data often suffers from high dimensionality, strong inter-protein correlations, feature redundancy, and missing values. Without effective data quality control, missing value handling, standardization, and feature screening mechanisms, it can easily lead to insufficient model stability, large prediction biases, and difficulty in identifying key aging biomarkers.
[0005] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0006] The main objective of this application is to propose a method and system for assessing extracellular matrix aging based on plasma proteomics, so as to achieve quantitative assessment of extracellular matrix aging status and improve the accuracy, stability and individualized characterization ability of extracellular mechanism aging assessment results.
[0007] To achieve the above objectives, one aspect of this application proposes a method for assessing extracellular matrix aging based on plasma proteomics, the method comprising the following steps: Obtain population sample data and plasma proteomics data to be tested; The population sample data is preprocessed to obtain the target population dataset; Based on a preset list of extracellular matrix proteins, extracellular matrix candidate protein data are extracted from the plasma proteomics data; The extracellular matrix candidate protein data were subjected to quality control, missing value imputation, and expression level standardization to obtain standardized protein expression data. Based on the target population dataset and the standardized protein expression data, an extracellular matrix aging assessment model was constructed to obtain the predicted age of the extracellular matrix and evaluate the model performance. Based on the deviation correction between the predicted age and the actual age of the extracellular matrix, the aging age difference of the extracellular matrix is obtained. Based on the predicted age of the extracellular matrix, the difference in aging age of the extracellular matrix, and the model performance indicators, an assessment result of extracellular matrix aging is generated.
[0008] In some embodiments, the plasma proteomics data includes sample identification data, protein identification data, protein expression level data, and protein loss rate data.
[0009] In some embodiments, the extraction of extracellular matrix candidate protein data from the plasma proteomics data based on a preset extracellular matrix protein list includes: Based on the functional categories of extracellular matrix proteins, a preset list of extracellular matrix proteins is determined; the preset list of extracellular matrix proteins includes elastic fibrous-related proteins, matrix metalloproteinase family proteins, collagen family proteins, and other extracellular matrix-related proteins. The preset list of extracellular matrix proteins is matched with the plasma proteomics data to obtain initial matching protein data; Extract the protein expression level data and protein loss rate data corresponding to the initial matched protein data to generate extracellular matrix candidate protein data.
[0010] In some embodiments, the process of performing quality control, missing value imputation, and expression level standardization on the extracellular matrix candidate protein data to obtain standardized protein expression data includes: The extracellular matrix candidate protein data were subjected to outlier removal and sample matching to obtain initial protein data; The initial protein data is filtered based on a preset missing rate threshold to obtain quality control protein data; For subsamples without any missing proteins, random missing values are created to verify the accuracy of imputation and evaluate the accuracy of multiple imputation. The missing values in the quality control protein data were filled using multiple interpolation to obtain complete protein data. The expression levels of the complete protein data were standardized to obtain standardized protein expression data.
[0011] In some embodiments, the step of constructing an extracellular matrix aging assessment model based on the target population dataset and the standardized protein expression data, obtaining the predicted age of the extracellular matrix, and evaluating the model performance includes: Based on the target population dataset and the standardized protein expression data, sample matching is performed to obtain the modeling dataset; The modeling dataset is divided according to a preset hierarchical rule to obtain a training dataset and a validation dataset; Based on the training dataset, an elastic network regression model was constructed and trained to obtain an extracellular matrix aging assessment model. The validation dataset is input into the extracellular matrix aging assessment model to obtain the predicted age of the extracellular matrix. The predictive performance of the extracellular matrix aging assessment model was evaluated using mean absolute error, and validated by comparing the difference between the predicted age and the actual age.
[0012] In some embodiments, the step of constructing and training a resilient network regression model based on the training dataset to obtain an extracellular matrix aging assessment model includes: A resilient network regression model is constructed using standardized protein expression data from the training dataset as input variables and corresponding actual age as output variables. The objective function of the elastic network regression model is constructed based on the prediction error term and regularization constraints; The regularization parameter in the objective function is optimized based on cross-validation and grid search to obtain the target regularization parameter; The elastic network regression model is trained based on the target regularization parameter to obtain the trained aging assessment model. Based on the trained aging assessment model, extracellular matrix aging-related proteins are identified, and an extracellular matrix aging assessment model is generated. Based on the extracellular matrix aging assessment model, an extracellular mechanism-predicted age is generated. The predictive performance of the extracellular matrix aging assessment model is evaluated using mean absolute error, and the model is validated by comparing the difference between the predicted age and the actual age.
[0013] In some embodiments, the objective function is formulated as follows: ; in, Represented by the vector of regression coefficients To optimize the object, find the model parameters that minimize the objective function; The output variable represents age; Represents a standardized protein expression matrix; Represents the regression coefficient vector; This represents the squared error term between the predicted age and the actual age. Indicates the penalty coefficient; This represents the ratio of the first regularization constraint; This represents the first norm of the regression coefficient vector; This represents the squared second norm of the regression coefficient vector.
[0014] In some embodiments, the step of correcting for the deviation between the predicted age and the actual age based on the extracellular matrix to obtain the extracellular matrix aging age difference includes: Based on the predicted age from the extracellular matrix and the actual age, sample pairing is performed to obtain age-paired data; Based on the age-paired data, a locally weighted scatter plot smoothing model is constructed to determine the baseline predicted age of the population corresponding to different actual ages. The initial aging age difference is obtained by calculating the difference between the extracellular matrix predicted age and the population baseline predicted age. The initial aging age difference is oriented and numerically corrected to obtain the extracellular matrix aging age difference.
[0015] In some embodiments, the calculation formula for the locally weighted scatter plot smoothing model is as follows: ; in, The locally weighted scatter plot smoothing method is represented in the first... Individual's actual age The fitted predicted age value obtained at this point is the baseline predicted age of the population; This indicates a local weighted scatter plot smoothing method; Indicates the first The actual age of each individual; Indicates the first The predicted age of an individual is output by the elastic network regression model; This indicates the proportion of neighboring data points used in local regression.
[0016] To achieve the above objectives, another aspect of this application proposes an extracellular matrix aging assessment system based on plasma proteomics, the system comprising: The data acquisition module is used to acquire population sample data and plasma proteomics data to be tested. The population data processing module is used to preprocess the population sample data to obtain the target population dataset; The candidate protein extraction module is used to extract extracellular matrix candidate protein data from the plasma proteomics data based on a preset extracellular matrix protein list. The protein data processing module is used to perform quality control, missing value imputation, and expression level standardization on the extracellular matrix candidate protein data to obtain standardized protein expression data. The evaluation model construction module is used to construct an extracellular matrix aging evaluation model based on the target population dataset and the standardized protein expression data, obtain the predicted age of the extracellular matrix and evaluate the model performance. The age difference calculation module is used to correct the deviation between the predicted age and the actual age based on the extracellular matrix to obtain the extracellular matrix aging age difference. The result generation module is used to generate extracellular matrix aging assessment results based on the predicted age of the extracellular matrix, the aging age difference of the extracellular matrix, and model performance indicators.
[0017] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0018] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0019] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0020] The embodiments of this application include at least the following beneficial effects: This application provides a method and system for assessing extracellular matrix aging based on plasma proteomics. This method improves data integrity, consistency, and comparability by acquiring and preprocessing population sample data and plasma proteomics data; by extracting candidate protein data based on a preset extracellular matrix protein list, the model construction process focuses on protein features related to extracellular matrix aging, reducing interference from irrelevant proteins; by performing quality control, missing value imputation, and expression level standardization on candidate protein data, the impact of outliers and missing data on the assessment results can be reduced; by constructing an extracellular matrix aging assessment model, it is possible to extract data reflecting cellular aging from plasma protein expression levels. The characteristic information of extracellular matrix aging status is used to obtain the predicted age of extracellular matrix, realizing the quantitative characterization of the degree of extracellular matrix aging. The predictive accuracy and generalization ability are verified by model performance to ensure that the assessment results have quantifiable confidence. By correcting the deviation between the predicted age and the actual age of extracellular matrix, the influence of age-related systematic bias on the assessment results can be reduced, and the extracellular matrix aging age difference can be obtained to better reflect the rate of individual aging. At the same time, the predicted age, extracellular matrix aging age difference and model performance indicators are combined to generate a comprehensive assessment result, making the assessment results individualized, quantifiable and verifiable, thereby improving the accuracy, stability and application value of extracellular matrix aging assessment. Attached Figure Description
[0021] Figure 1 This is a schematic flowchart of an extracellular matrix aging assessment method based on plasma proteomics provided in an embodiment of this application; Figure 2 This is a graph showing the regression coefficients of key proteins in the male extracellular matrix aging assessment model provided in this application embodiment; Figure 3 This is a graph showing the regression coefficients of key proteins in the extracellular matrix aging assessment model for female populations provided in this application embodiment; Figure 4 This is a combination of density heat map, density distribution and LOWESS fitting curve of extracellular matrix predicted age and actual age of male population provided in the embodiments of this application; Figure 5 This is a combination of density heat map, density distribution and LOWESS fitting curve of extracellular matrix predicted age and actual age of female population provided in the embodiments of this application; Figure 6 This is a schematic diagram of a module of an extracellular matrix aging assessment system based on plasma proteomics provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0023] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0024] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0026] Before providing a detailed description of the embodiments of this application, some of the nouns and terms used in the embodiments of this application will be explained first. The nouns and terms used in the embodiments of this application shall be interpreted as follows: Plasma proteomics refers to a data analysis technique that uses proteins in plasma samples as the detection targets and obtains information on the expression levels, detection quality, and functional annotation of different proteins through high-throughput detection technology. It is used to reflect the physiological state, pathological changes, and aging-related molecular characteristics of the body. The extracellular matrix (ECM) is a complex network structure located outside cells, composed of components such as collagen, elastin, glycoproteins, and proteoglycans. It is used to maintain tissue structure stability, mechanical properties, and cellular microenvironment homeostasis. Extracellular matrix proteins are proteins that participate in the composition, remodeling, degradation, or functional regulation of the extracellular matrix, including elastin-associated proteins, matrix metalloproteinase family proteins, collagen family proteins, and other extracellular matrix-associated proteins. Elastic fiber-related proteins are proteins involved in the formation, maintenance, and degradation of elastic fibers, and are used to characterize the integrity of tissue elastic structure and the degenerative changes of elastic fibers during aging. Matrix metalloproteinases (MMPs) are a family of proteases that can degrade components of the extracellular matrix and are used to reflect changes related to extracellular matrix remodeling, degradation, and tissue aging. Collagen family proteins are a class of proteins that constitute the main structural framework of the extracellular matrix. Their composition ratio, cross-linking state, and changes in the dynamic balance of synthesis and degradation are used to characterize the maintenance and degenerative changes of tissue strength, structural integrity, and mechanical properties. A UK Biobank is a large, prospective population cohort database that contains information on demographic characteristics, health status, disease outcomes, and multi-omics data, which can be used for disease risk analysis, aging assessment, and biomarker research. Chronic obstructive pulmonary disease (COPD) is a chronic respiratory disease characterized by persistent airflow limitation. Its development is related to extracellular matrix remodeling and tissue degenerative changes. Protein expression data refers to data used to characterize the detection level of target proteins in plasma samples, and can be used as input features for constructing extracellular matrix aging assessment models; Protein missing rate data refers to the proportion of target proteins that were not effectively detected or whose expression levels were not recorded in the sample population. It is used to determine whether the protein data meets the requirements for subsequent modeling and analysis. Elastic Net Regression Model (ENRM) is a regression model that combines sparse constraints and coefficient shrinkage constraints. It is used to screen key proteins related to extracellular matrix aging in high-dimensional protein expression data and to build an age prediction model. Locally Weighted Scatterplot Smoothing (LOESS) is a smoothing model used to perform curve fitting by weighting local neighbor samples with a preset proportion of neighboring data points. It is used to determine the baseline predicted age of different actual ages.
[0027] This application provides a method and system for assessing extracellular matrix aging based on plasma proteomics. This approach improves data integrity, consistency, and comparability by acquiring and preprocessing population sample data and plasma proteomics data. By extracting candidate protein data based on a pre-defined extracellular matrix protein list, the model construction process focuses on protein features related to extracellular matrix aging, reducing interference from irrelevant proteins. Quality control, missing value imputation, and expression level standardization of candidate protein data reduce the impact of outliers and missing data on the assessment results. By constructing an extracellular matrix aging assessment model, it is possible to extract data reflecting the aging status of the extracellular matrix from plasma protein expression levels. By analyzing the characteristic information of the extracellular matrix, the predicted age of the extracellular matrix is obtained, enabling a quantitative characterization of the degree of extracellular matrix aging. The predictive accuracy and generalization ability are verified through model performance, ensuring that the assessment results have quantifiable confidence. By correcting the deviation between the predicted age and the actual age of the extracellular matrix, the influence of age-related systematic bias on the assessment results can be reduced, further obtaining the extracellular matrix aging age difference, which better reflects the rate of individual aging. At the same time, by combining the predicted age of the extracellular matrix, the extracellular matrix aging age difference, and model performance indicators, a comprehensive assessment result is generated, making the assessment results individualized, quantifiable, and verifiable, thereby improving the accuracy, stability, and application value of extracellular matrix aging assessment.
[0028] This application provides a method for assessing extracellular matrix aging based on plasma proteomics, relating to the field of aging assessment technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited thereto. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the plasma proteomics-based extracellular matrix aging assessment method, but is not limited to the above forms.
[0029] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0030] Figure 1 This is an optional flowchart of an extracellular matrix aging assessment method based on plasma proteomics provided in this application embodiment. Figure 1 The method may include, but is not limited to, steps S1 to S7: S1: Obtain the sample data of the population to be tested and the plasma proteomics data; the plasma proteomics data includes sample identification data, protein identification data, protein expression level data and protein loss rate data.
[0031] In this embodiment, population sample data of the test subjects are extracted from the prospective cohort study dataset of the UK Biobank according to a pre-defined list of variables. This includes sample number, actual age, sex, baseline health status, and disease status information related to extracellular matrix aging. During data extraction, the sample identifier is used as the unique matching field, and population information from different sources is uniformly encoded and structured for storage to ensure accurate subsequent association with proteomics data.
[0032] Simultaneously, plasma proteomics data for the corresponding samples were acquired. This plasma proteomics data includes sample identification data, protein identification data, protein expression level data, and protein deletion rate data. The sample identification data is used to correlate with population sample data; the protein identification data is used to distinguish different proteins; the protein expression level data is used to characterize the detection level of each protein in the plasma sample; and the protein deletion rate data is used to determine the reliability of the protein data.
[0033] After data acquisition, a consistency check was performed on the population sample data and plasma proteomics data to determine if there were any duplicates, missing, or mismatched sample identifiers, and data records that could not be correlated were removed. Subsequently, the population sample data and protein expression data were merged according to the sample identifiers to form a raw analysis dataset containing sample information, actual age, protein expression level, and protein missing information, providing basic data for subsequent population screening, candidate protein extraction, and model construction.
[0034] S2: Preprocess the population sample data to obtain the target population dataset; In this embodiment, the population sample data is cleaned and filtered. Specifically, the actual age and gender of the samples are first checked for missing, abnormal, or logically conflicting information, such as actual age exceeding a preset range, gender field being unidentifiable, duplicate records of the same sample, or incomplete disease status information. Abnormal samples that cannot be corrected are removed, and data fields with inconsistent formats are uniformly converted.
[0035] The target study population was determined based on pre-defined inclusion and exclusion criteria. Inclusion criteria included complete sample identification, valid actual age information, matching plasma proteomics data, and a complete health status record. Exclusion criteria included baseline diagnosis of diseases associated with severe extracellular matrix aging (such as chronic obstructive pulmonary disease, arteriosclerosis, osteoarthritis, etc.), excessively high proportion of missing key variables, protein data quality not meeting requirements, or sample identification not being uniquely matched. Through these screenings, the target population dataset met the requirements for subsequent statistical analysis and model training. In this embodiment, 530,111 eligible participants were ultimately included.
[0036] After sample selection, the retained population variables were normalized. Categorical variables such as gender were coded, while continuous variables such as age underwent statistical description, outlier identification, and necessary scaling transformations. The final target population dataset includes at least sample identifiers, actual age, gender, and the population characteristic fields required for modeling, and maintains a correspondence with the subsequent standardized protein expression data at the sample level.
[0037] Specifically, a systematic statistical description is performed on the baseline characteristics of the test population and the distribution of various variables to comprehensively characterize the sample. For continuous variables, appropriate statistical indicators are selected for description based on their distribution characteristics (such as whether they follow a normal distribution), including mean ± standard deviation or median (interquartile range). For categorical variables, frequency and composition ratio are used for description.
[0038] Systematic standardization and preprocessing operations were performed on all types of variables included in the analysis to improve data consistency and comparability, and to meet the input requirements of statistical models and machine learning algorithms. Specifically, this included dummy variable encoding for categorical variables and normalization or centering for continuous variables; logarithmic transformation was performed on variables with significant distribution skewness to improve their distribution characteristics; and Z-score standardization and other methods were used to unify the scale of variables, thereby reducing the impact of differences in units on the model results.
[0039] S3: Extract extracellular matrix candidate protein data from plasma proteomics data based on a pre-defined list of extracellular matrix proteins; Specifically, based on a pre-defined list of extracellular matrix proteins, candidate extracellular matrix proteins are extracted from plasma proteomics data, including: Based on the functional categories of extracellular matrix proteins, a preset list of extracellular matrix proteins is determined; this preset list of extracellular matrix proteins includes elastic fibrous-related proteins, matrix metalloproteinase family proteins, collagen family proteins, and other extracellular matrix-related proteins. The pre-defined list of extracellular matrix proteins was matched with plasma proteomics data to obtain initial matching protein data; Extract protein expression level data and protein loss rate data corresponding to the initial matching protein data to generate extracellular matrix candidate protein data.
[0040] In this embodiment, a preset list of extracellular matrix proteins is determined based on their functional properties. Specifically, proteins involved in maintaining extracellular matrix structure, elastic fiber formation, collagen deposition, matrix degradation, and tissue remodeling are included in the screening scope, including elastic fiber-related proteins (FBLN2, ELN, FBN2), matrix metalloproteinase family proteins (MMP1, MMP3, MMP7, MMP8, MMP9, MMP10, MMP12, MMP13, MMP15), collagen family proteins (COL1A1, COL5A1, COL3A1, COL4A1), and other extracellular matrix-related proteins (SPARC, THBS2, SFRP4, TNC, FN1, LAMB1).
[0041] After obtaining a pre-defined list of extracellular matrix proteins, this list is matched against plasma proteomics data. During matching, protein identification data is the primary basis, supplemented by protein name, protein code, and protein family affiliation for auxiliary identification, reducing missed matches due to differences in naming, abbreviations, or detection platform fields. For cases where the same protein corresponds to multiple detection items, initial matching protein data can be obtained by filtering based on the completeness of protein detection annotations and sample coverage.
[0042] Protein expression levels and protein loss rates were extracted from the initial matched protein data and organized into extracellular matrix candidate protein data according to sample and protein identifiers. These data were then used for subsequent quality control, missing value imputation, expression level standardization, and aging assessment model construction.
[0043] S4: Perform quality control, missing value imputation, and expression level standardization on extracellular matrix candidate protein data to obtain standardized protein expression data; The extracellular matrix candidate protein data underwent quality control, missing value imputation, and expression level standardization to obtain standardized protein expression data, including: Outlier removal and sample matching were performed on the extracellular matrix candidate protein data to obtain initial protein data; The initial protein data is filtered based on a preset missing rate threshold to obtain quality control protein data; For subsamples without any missing proteins, random missing values are created to verify the accuracy of imputation and evaluate the accuracy of multiple imputation. Multiple imputation was used to fill in missing values in the quality control protein data to obtain complete protein data. The expression levels of the complete protein data were standardized to obtain standardized protein expression data.
[0044] In this embodiment, outlier removal and sample matching are performed on the extracellular matrix candidate protein data. Specifically, the candidate protein expression data are matched with the target population dataset based on the sample identifier, and records with missing, duplicate, or mismatched sample identifiers are removed. At the same time, based on the protein detection quality data and expression level distribution, data records with abnormal detection quality, expression levels that deviate significantly from the overall distribution, or that do not meet the detection requirements are removed, resulting in initial protein data where the samples and protein expression information correspond consistently.
[0045] The initial protein data is screened based on a preset missing rate threshold. The missing percentage of each candidate protein in the target sample is calculated, and proteins with a missing rate exceeding the preset threshold are removed. For example, proteins with a missing value exceeding 10% are removed to ensure good data integrity and stability for proteins entering subsequent analysis.
[0046] To verify the accuracy of imputation, missing values are randomly created in a subset of protein data that has no missing proteins. Specifically, missing values are randomly set in the subset of protein data without missing values, and the imputed values are compared with the original true values. If the imputation deviation is within an acceptable range, the multiple imputation method is deemed suitable; if the deviation is large, the imputation method is adjusted. For example, if the overall missing value rate of the original protein dataset is 3%, missing values are randomly created in a subset of protein data without missing values as a validation set, and the imputation accuracy is evaluated by comparing the difference between the imputed values and the true values.
[0047] After successful validation, for proteins with a missing rate not exceeding the threshold, multiple imputation was used to fill in the remaining missing values. This involved multiple estimations based on observed protein expression levels and inter-protein correlations, followed by merging the imputation results to obtain complete protein data. Protein expression levels were then processed using logarithmic transformation, centering, normalization, or standard deviation scaling to ensure that different protein expression levels were on a uniform scale, ultimately yielding standardized protein expression data for subsequent training and prediction of extracellular matrix aging assessment models.
[0048] Specifically, outlier detection was performed on the extracellular matrix candidate protein data to remove abnormal data; the cleaned protein expression data was matched and horizontally merged with the population baseline information (such as age, gender, etc.) according to the sample ID to form a complete analysis dataset; Calculate the percentage of missing values for each protein and remove proteins with more than 10% missing values (9 proteins) to ensure that the proteins included in the analysis have complete data quality; In this embodiment, the overall missing value rate of the original protein dataset is 3%. To evaluate the imputation effect, missing values were randomly created artificially as a validation set in a subset of protein data with no missing values. The accuracy of imputation was evaluated by comparing the difference between the imputed estimated values and the true values. The mean absolute error of the imputation results at the original scale was verified to be 0.591. The formula for calculating the mean absolute error is as follows: ; in, To verify the sample size in the set. Let i be the true value of the i-th sample. This is the estimated value for this sample obtained by the interpolation method. The mean absolute error is denoted as 0. The closer the mean absolute error is to 0, the smaller the average deviation between the interpolated estimated value and the true value, and the more ideal the interpolation effect.
[0049] After the validation is passed, the remaining missing values of the proteins retained after the above screening (missing values ≤10%) are filled with multiple imputation to obtain a complete and reliable protein expression dataset.
[0050] Multiple interpolation follows Rubin's Rules, and its core process is to generate multiple complete datasets, analyze each dataset separately, and finally merge the parameter estimation results of each dataset.
[0051] Protein data are normalized or centered; logarithmic transformation is performed on variables with significant skewness to improve their distribution characteristics; and Z-score standardization and other methods are used to unify the scale of variables, thereby reducing the impact of dimensional differences on model results.
[0052] S5: Construct an extracellular matrix aging assessment model based on the target population dataset and standardized protein expression data, obtain the predicted age of the extracellular matrix and evaluate the model performance; This includes constructing an extracellular matrix aging assessment model based on the target population dataset and standardized protein expression data, obtaining the predicted age of the extracellular matrix, and evaluating the model performance, including: The modeling dataset is obtained by matching samples between the target population dataset and standardized protein expression data; The modeling dataset is divided according to a preset hierarchical rule to obtain a training dataset and a validation dataset; An elastic network regression model was constructed and trained based on the training dataset to obtain an extracellular matrix aging assessment model; Input the validation dataset into the extracellular matrix aging assessment model to obtain the predicted age of the extracellular matrix; The predictive performance of the extracellular matrix aging assessment model was evaluated using mean absolute error (MAE), and validated by comparing the difference between predicted age and actual age.
[0053] In this embodiment, the target population dataset and standardized protein expression data are merged using sample identifiers as the matching field to obtain the modeling dataset. Specifically, information such as actual age, gender, and baseline health status in the target population dataset are integrated with the standardized extracellular matrix protein expression levels of the corresponding samples. Records with missing, duplicate, or unmatched sample identifiers, or incomplete protein expression data, are removed or marked to ensure that each sample in the modeling dataset has complete actual age information and protein expression characteristics.
[0054] The modeling dataset is divided according to a pre-defined stratification rule to obtain a training dataset and a validation dataset. Specifically, the stratification can be based on the age distribution of the participants to ensure consistency in age structure between the training and validation datasets, reducing bias caused by data partitioning. The dataset is divided in a 7:3 ratio, with the training dataset used for model parameter estimation, key protein screening, and model training, and the validation dataset used to test the model's predictive ability on samples not used in training. Male and female samples can be partitioned and modeled separately to improve the model's adaptability to different gender groups.
[0055] An elastic network regression model was constructed and trained based on the training dataset to obtain an extracellular matrix aging assessment model. Specifically, the standardized extracellular matrix protein expression levels in the training dataset were used as input variables, and actual age was used as the output variable. Model parameters were determined through cross-validation and parameter search to enable the model to identify key proteins related to extracellular matrix aging and reduce the impact of redundant protein variables on the prediction results. After model training, the standardized protein expression levels in the validation dataset were input into the extracellular matrix aging assessment model, which outputs a continuous predicted extracellular matrix age for each sample. The predicted age was then paired and saved with sample identifier, actual age, and gender information for subsequent model performance evaluation and calculation of aging age differences.
[0056] Specifically, an elastic network regression model is constructed and trained based on the training dataset to obtain an extracellular matrix aging assessment model, including: A resilient network regression model was constructed using standardized protein expression data from the training dataset as input variables and corresponding actual age as output variables. The objective function of the elastic network regression model is constructed based on the prediction error term and regularization constraints; The regularization parameter in the objective function is optimized based on cross-validation and grid search to obtain the target regularization parameter; The elastic network regression model is trained based on the target regularization parameter to obtain the trained aging assessment model. Based on the trained aging assessment model, extracellular matrix aging-related proteins are identified, and an extracellular matrix aging assessment model is generated. Based on the extracellular matrix aging assessment model, an extracellular mechanism is used to predict age. The predictive performance of the extracellular matrix aging assessment model is evaluated by mean absolute error, and the model is validated by comparing the difference between the predicted age and the actual age.
[0057] In this embodiment, a resilient network regression model is constructed using standardized protein expression data from the training dataset as input variables and the actual age of the corresponding samples as output variables. Specifically, the expression levels of candidate extracellular matrix proteins for each sample are organized into a model input matrix, and the actual age is used as the model learning objective, enabling the model to learn the correspondence between protein expression levels and age changes. Simultaneously, a model objective function is constructed based on a prediction error term and regularization constraints. The prediction error term measures the deviation between the predicted age and the actual age, while the regularization penalty term consists of L1 and L2 regularization. L1 regularization is used to screen key protein features, and L2 regularization is used to handle collinearity among proteins and compress the influence of redundant variables, thereby improving model stability while ensuring prediction accuracy.
[0058] Subsequently, the regularization parameter in the objective function is optimized based on cross-validation and grid search. Specifically, the training dataset is divided into multiple subsets, and model training and validation are performed iteratively. Prediction error and model stability are compared under different parameter combinations. A combination of 10-fold cross-validation and grid search is used to select the parameter combination with smaller prediction error and better stability as the target regularization parameter, thereby determining the feature selection strength and overall constraint degree of the model.
[0059] After determining the target regularization parameter, the elastic network regression model is retrained based on this parameter to obtain a trained aging assessment model. After training, protein variables with non-zero regression coefficients or contributions meeting preset requirements are extracted from the model as extracellular matrix aging-related proteins; proteins with compressed coefficients or weak contributions are not included as core features. This generates an extracellular matrix aging assessment model containing key protein features, model parameters, and prediction rules, which is used to subsequently output predicted extracellular matrix age based on standardized protein expression data.
[0060] Specifically, standardized protein expression levels were used as the predictor variable, and age as the response variable. A 10-fold cross-validation method was used to train the elastic regression model on the training set. The training set was randomly divided into 10 subsets, with 9 subsets used for model training and 1 subset used for validation, repeated 10 times. A grid search strategy was used to systematically optimize the regularization hyperparameters (L1 ratio, penalty coefficient λ). The optimal parameter combination was selected based on minimizing the mean squared error of cross-validation, and the λ value corresponding to the minimum mean squared error was recorded as the regularization parameter for the final model. The objective function of the elastic regression model (loss function + regularization term) is as follows: ; in, Represented by the vector of regression coefficients To optimize the object, find the model parameters that minimize the objective function; The output variable represents age; Represents a standardized protein expression matrix; Represents the regression coefficient vector; This represents the squared error term between the predicted age and the actual age. Indicates the penalty coefficient; This represents the ratio of the first regularization constraint; This represents the first norm of the regression coefficient vector; This represents the squared second norm of the regression coefficient vector.
[0061] Based on the optimal parameters determined by the objective function, the elastic network regression model is refitted. Utilizing the L1 regularization property of the elastic network, the regression coefficients of redundant proteins are compressed to zero, thereby automatically performing feature selection while estimating parameters and effectively handling the multicollinearity problem among proteins. Protein variables with non-zero regression coefficients are extracted to constitute a subset of key proteins significantly associated with extracellular matrix (ECM) aging, i.e., the core protein set of the ECM aging index. In this embodiment, the selected subset of key proteins differs between men and women. Key proteins selected from the male population include: MMP1, MMP3, MMP7, MMP8, MMP9, MMP10, MMP12, MMP13, COL1A1, COL4A1, SPARC, THBS2, and TNC, totaling 13. Key proteins selected from the female population include: MMP1, MMP3, MMP7, MMP8, MMP10, MMP12, MMP13, COL1A1, COL4A1, SPARC, THBS2, and TNC, totaling 12. These protein sets serve as the core proteins for the male and female ECM aging indices, respectively, and will be used in subsequent age prediction model construction. In the healthy subgroup without baseline ECM-related diseases, the model training and variable selection process was repeated for each gender to construct elastic network regression models and ECM aging index core protein sets for men and women in the healthy population. These models were then used for comparative analysis with the full population model.
[0062] Using a trained elastic regression network model, age prediction is performed on the corresponding samples. The standardized protein expression levels of each individual are input into the model. The model calculates the predicted age for each individual based on the key protein subsets selected by each model and their corresponding regression coefficients. The predicted age is output as a continuous variable and paired with the individual's actual age for storage, forming a structured dataset containing individual ID, actual age, and predicted age. This dataset is used for subsequent model performance evaluation and ECM aging age difference calculation steps.
[0063] refer to Figure 2 As shown, Figure 2 The regression coefficients of key proteins selected by the elastic network regression model in the male population are plotted to show the direction and magnitude of each protein's contribution to age prediction. Figure 2 The horizontal axis represents the protein name, and the vertical axis represents the regression coefficient. A positive coefficient indicates that an increase in the expression level of the corresponding protein has a positive contribution to the prediction of age of the extracellular matrix, while a negative coefficient indicates that an increase in the expression level of the corresponding protein has a negative contribution to the prediction of age of the extracellular matrix. Figure 2 The model coefficients of MMP family proteins, collagen family proteins, and other extracellular matrix-related proteins obtained from screening corresponding male models are displayed.
[0064] from Figure 2 As can be seen, the coefficients of MMP12, MMP7, COL4A1, MMP1, MMP3, MMP9, MMP13, and TNC are positive, with MMP12 having the largest positive coefficient, indicating that it has the highest weight in the calculation of male extracellular matrix-related age. The coefficients of MMP8, THBS2, MMP10, SPARC, and COL1A1 are negative, indicating that these proteins have a negative contribution to the predicted age in the model. Overall, this figure intuitively reflects the differences in the weights of different extracellular matrix-related proteins in the male aging assessment model, and can serve as an important basis for interpreting extracellular matrix-related age prediction and subsequent calculation of age differences.
[0065] refer to Figure 3 As shown, Figure 3 This is a regression coefficient diagram of key proteins in an extracellular matrix aging assessment model for a female population. It is used to show the extracellular matrix-related proteins screened by the elastic network regression model and their contribution direction and magnitude to age prediction. Figure 3 Among them, the coefficients of MMP12, MMP7, MMP13, MMP3, COL1A1, TNC, and MMP1 are positive, indicating that their expression changes have a positive contribution to the prediction of age in female extracellular matrix. MMP12 has the highest weight, indicating that it has a larger weight in the female model. The coefficients of SPARC, MMP8, COL4A1, THBS2, and MMP10 are negative, indicating that they have a negative contribution to the prediction of age in the model. Overall, Figure 3 This reflects the differences in the contributions of different extracellular matrix-related proteins in aging assessment models within the female population. It can be used to explain the basis for extracellular matrix-predicted age and provide a model feature foundation for subsequent calculation of extracellular matrix aging age differences.
[0066] refer to Figure 4 As shown, Figure 4 This is a composite graph of predicted and actual ages of the extracellular matrix in a male population, used to illustrate the predicted distribution and fitting effect of the extracellular matrix aging assessment model. Figure 4 The main plot at the bottom center uses actual age as the horizontal axis and predicted age as the vertical axis. The scatter plot and color density represent the concentration of samples in different age groups. The black trend line reflects that the predicted age increases as the actual age increases, indicating that the model can capture the relationship between the expression of extracellular matrix-related proteins and age changes. The density curve on the right shows the overall distribution of predicted age, and the bar chart at the top shows the distribution of the number of samples in different actual age groups. Figure 4 The mean absolute error of the model is 6.21 years, indicating that there is a certain deviation between the model's predicted age and the actual age. This error level is within the acceptable range of model performance. Figure 4The corresponding model performance evaluation and prediction result visualization stage can provide a data foundation for subsequent deviation correction based on predicted age and actual age, and for calculating the age difference of extracellular matrix aging.
[0067] refer to Figure 5 As shown, Figure 5 This is a composite graph showing the predicted age and actual age of the extracellular matrix in a female population, used to demonstrate the predictive effectiveness and sample distribution of the extracellular matrix aging assessment model. Figure 5 The main chart at the bottom center uses actual age as the horizontal axis and predicted age as the vertical axis. The color density represents the degree of sample clustering. The black trend line shows that the predicted age increases overall with the increase of actual age, indicating that the model can capture the aging characteristics related to age changes in the female population based on the expression level of plasma extracellular matrix-related proteins. The bar chart at the top reflects the distribution of sample size in different actual age groups, and the density curve on the right reflects the overall distribution of predicted age. Figure 5 The mean absolute error (MAE) is 5.9 years, indicating that there is a certain deviation between the predicted age and the actual age of the female model. This error level is within the acceptable range of model performance. Figure 5 It can be used to illustrate the validation results of the extracellular matrix aging assessment model and provide a basis for subsequent deviation correction based on predicted age and actual age and calculation of the extracellular matrix aging age difference.
[0068] S6: Based on the deviation correction between the predicted age and the actual age of the extracellular matrix, the aging age difference of the extracellular matrix is obtained; Among them, the extracellular matrix aging age difference is obtained by correcting for the discrepancy between the predicted age and the actual age based on the extracellular matrix, including: Age-paired data is obtained by pairing samples based on predicted age from extracellular matrix with actual age. A locally weighted scatter plot smoothing model is constructed based on age-paired data to determine the baseline predicted age of the population corresponding to different actual ages. The initial aging age difference is obtained by calculating the difference between the predicted age based on the extracellular matrix and the population baseline predicted age. The initial aging age difference was oriented and numerically corrected to obtain the extracellular matrix aging age difference.
[0069] In this embodiment, sample pairing is performed based on the predicted age from the extracellular matrix and the actual age to obtain age-paired data. Specifically, using the sample identifier as the association field, the predicted age from the extracellular matrix output by the model is merged with the actual age in the target population dataset, while simultaneously retaining information such as gender, sample source, and model version. Records with missing predicted age, missing actual age, duplicate sample identifiers, or unmatched records are removed or marked to ensure that each sample has complete and consistent age-paired information.
[0070] A locally weighted scatter plot smoothing model is constructed based on age-paired data. Specifically, using actual age as the independent variable and predicted age from the extracellular matrix as the dependent variable, a locally smoothed fit is performed on the predicted age at different actual age levels. During the fitting process, for each actual age or adjacent age interval, the predicted age data of neighboring samples are selected and weighted to obtain the average predicted age of the population at that age level.
[0071] After obtaining the smoothed fitting results, the fitted predicted values corresponding to different actual ages are determined as the baseline predicted age for the population. This baseline predicted age is used to characterize the average prediction level of the same-age population in the extracellular matrix aging assessment model, and can be used to correct for systematic biases between predicted age and actual age in different age groups. This process avoids age group bias caused by directly subtracting the predicted age from the actual age, making subsequent age differences more reflective of the degree of aging deviation of an individual relative to their peers.
[0072] Finally, the initial aging age difference is calculated by comparing the predicted extracellular matrix age of each sample with the corresponding actual age of the population baseline. This initial aging age difference is then oriented and numerically corrected. Specifically, a difference greater than zero is categorized as relatively accelerated extracellular matrix aging, a difference less than zero is categorized as relatively delayed extracellular matrix aging, and a difference close to zero is categorized as close to the average level of the same age group. Furthermore, abnormally large differences can be validated or robustned to ultimately obtain the extracellular matrix aging age difference.
[0073] Specifically, the calculation formula for the locally weighted scatter plot smoothing model is as follows: ; in, The locally weighted scatter plot smoothing method is represented in the first... Individual's actual age The fitted predicted age value obtained at this point is the baseline predicted age of the population; This indicates a local weighted scatter plot smoothing method; Indicates the first The actual age of each individual; Indicates the first The predicted age of an individual is output by the elastic network regression model; Indicates the proportion of neighboring data points used in local regression; This indicates the proportion of data points used in local regression. For each target point, approximately two-thirds of its neighboring sample points are selected for local weighted regression to control the smoothness of the fitted curve. A larger span value results in a smoother fitted curve; a smaller span value results in a curve that fits the local data more closely.
[0074] Calculate the age difference for each individual, which is the individual's predicted age minus the population's average predicted age at that individual's actual age, estimated by the locally weighted scatter plot smoothing model:
[0075] in, Let be the extracellular matrix aging age difference of the i-th individual. >0 indicates that the individual's predicted age is higher than the average for their age, suggesting accelerated ECM aging; A value less than 0 indicates that the level is below the average for the same age group, suggesting delayed ECM aging.
[0076] S7: Generate extracellular matrix aging assessment results based on extracellular matrix predicted age, extracellular matrix aging age difference, and model performance indicators.
[0077] In this embodiment, model performance is evaluated using metrics including the coefficient of determination, mean absolute error (MAO), and the performance difference between the training and validation sets. The MAO reflects the consistency between the model's predicted age and the actual age, the MAO reflects the magnitude of the prediction error, and the difference between the training and validation set metrics is used to determine whether the model is significantly overfitting. These metrics allow for a comprehensive assessment of the reliability and generalization ability of the extracellular matrix aging assessment model.
[0078] The predicted age, actual age, extracellular matrix aging age difference, and model performance indicators of each sample are integrated to generate individual-level extracellular matrix aging assessment results. The assessment results include sample identification, actual age, predicted age, aging age difference, aging status determination, and the model version used. Based on the sign and magnitude of the aging age difference, individuals can be classified into states such as accelerated extracellular matrix aging, near-average age for their peers, or delayed extracellular matrix aging.
[0079] Finally, a structured assessment report is generated. This report includes candidate protein screening results, key protein subsets, model performance metrics, the relationship between predicted age and actual age, the distribution of extracellular matrix aging age differences, and individual assessment conclusions. For population studies, it can further output the distribution characteristics of extracellular matrix aging in different genders, age groups, or health statuses; for individual assessments, it can output the degree of deviation of an individual's extracellular matrix aging relative to their peers, providing data support for aging risk stratification, early screening of chronic diseases, and subsequent intervention studies.
[0080] Please see Figure 6 This application also provides an extracellular matrix aging assessment system based on plasma proteomics, the system comprising: The data acquisition module is used to acquire population sample data and plasma proteomics data to be tested. The population data processing module is used to preprocess population sample data to obtain the target population dataset; The candidate protein extraction module is used to extract extracellular matrix candidate protein data from plasma proteomics data based on a preset extracellular matrix protein list. The protein data processing module is used to perform quality control, missing value imputation, and expression level standardization on extracellular matrix candidate protein data to obtain standardized protein expression data. The evaluation model building module is used to construct an extracellular matrix aging assessment model based on the target population dataset and standardized protein expression data, obtain the predicted age of the extracellular matrix and evaluate the model performance. The age difference calculation module is used to correct the deviation between the predicted age and the actual age based on the extracellular matrix, and to obtain the extracellular matrix aging age difference. The results generation module is used to generate extracellular matrix aging assessment results based on extracellular matrix predicted age, extracellular matrix aging age difference, and model performance indicators.
[0081] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0082] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0083] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0084] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0085] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0086] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0087] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0088] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0089] This application provides a method and system for assessing extracellular matrix aging based on plasma proteomics. This approach improves data integrity, consistency, and comparability by acquiring and preprocessing population sample data and plasma proteomics data. By extracting candidate protein data based on a pre-defined extracellular matrix protein list, the model construction process focuses on protein features related to extracellular matrix aging, reducing interference from irrelevant proteins. Quality control, missing value imputation, and expression level standardization of candidate protein data reduce the impact of outliers and missing data on the assessment results. By constructing an extracellular matrix aging assessment model, it is possible to extract data reflecting the aging status of the extracellular matrix from plasma protein expression levels. By analyzing the characteristic information of the extracellular matrix, the predicted age of the extracellular matrix is obtained, enabling a quantitative characterization of the degree of extracellular matrix aging. The predictive accuracy and generalization ability are verified through model performance, ensuring that the assessment results have quantifiable confidence. By correcting the deviation between the predicted age and the actual age of the extracellular matrix, the influence of age-related systematic bias on the assessment results can be reduced, further obtaining the extracellular matrix aging age difference, which better reflects the rate of individual aging. At the same time, by combining the predicted age of the extracellular matrix, the extracellular matrix aging age difference, and model performance indicators, a comprehensive assessment result is generated, making the assessment results individualized, quantifiable, and verifiable, thereby improving the accuracy, stability, and application value of extracellular matrix aging assessment.
[0090] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0091] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0092] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0093] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0094] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for assessing extracellular matrix aging based on plasma proteomics, characterized in that, The method includes the following steps: Obtain population sample data and plasma proteomics data to be tested; The population sample data is preprocessed to obtain the target population dataset; Based on a pre-defined list of extracellular matrix proteins, extracellular matrix candidate protein data are extracted from the plasma proteomics data. The extracellular matrix candidate protein data were subjected to quality control, missing value imputation, and expression level standardization to obtain standardized protein expression data. Based on the target population dataset and the standardized protein expression data, an extracellular matrix aging assessment model was constructed to obtain the predicted age of the extracellular matrix and evaluate the model performance. Based on the deviation correction between the predicted age and the actual age of the extracellular matrix, the aging age difference of the extracellular matrix is obtained. Based on the predicted age of the extracellular matrix, the difference in aging age of the extracellular matrix, and the model performance indicators, an assessment result of extracellular matrix aging is generated.
2. The method according to claim 1, characterized in that, The plasma proteomics data includes sample identification data, protein identification data, protein expression level data, and protein loss rate data.
3. The method according to claim 1, characterized in that, The extraction of extracellular matrix candidate protein data from the plasma proteomics data based on a preset extracellular matrix protein list includes: Based on the functional categories of extracellular matrix proteins, a preset list of extracellular matrix proteins is determined; the preset list of extracellular matrix proteins includes elastic fibrous-related proteins, matrix metalloproteinase family proteins, collagen family proteins, and other extracellular matrix-related proteins. The preset list of extracellular matrix proteins is matched with the plasma proteomics data to obtain initial matching protein data; Extract the protein expression level data and protein loss rate data corresponding to the initial matched protein data to generate extracellular matrix candidate protein data.
4. The method according to claim 1, characterized in that, The process of quality control, missing value imputation, and expression level standardization of the extracellular matrix candidate protein data to obtain standardized protein expression data includes: The extracellular matrix candidate protein data were subjected to outlier removal and sample matching to obtain initial protein data; The initial protein data is filtered based on a preset missing rate threshold to obtain quality control protein data; For subsamples without any missing proteins, random missing values are created to verify the accuracy of imputation and evaluate the accuracy of multiple imputation. The missing values in the quality control protein data were filled using multiple interpolation to obtain complete protein data. The expression levels of the complete protein data were standardized to obtain standardized protein expression data.
5. The method according to claim 1, characterized in that, The process of constructing an extracellular matrix aging assessment model based on the target population dataset and the standardized protein expression data, obtaining the predicted age of the extracellular matrix, and evaluating the model performance includes: Based on the target population dataset and the standardized protein expression data, sample matching is performed to obtain the modeling dataset; The modeling dataset is divided according to a preset hierarchical rule to obtain a training dataset and a validation dataset; Based on the training dataset, an elastic network regression model was constructed and trained to obtain an extracellular matrix aging assessment model. The validation dataset is input into the extracellular matrix aging assessment model to obtain the predicted age of the extracellular matrix. The predictive performance of the extracellular matrix aging assessment model was evaluated using mean absolute error, and validated by comparing the difference between the predicted age and the actual age.
6. The method according to claim 5, characterized in that, The process of constructing and training a resilient network regression model based on the training dataset to obtain an extracellular matrix aging assessment model includes: A resilient network regression model is constructed using standardized protein expression data from the training dataset as input variables and corresponding actual age as output variables. The objective function of the elastic network regression model is constructed based on the prediction error term and regularization constraints; The regularization parameter in the objective function is optimized based on cross-validation and grid search to obtain the target regularization parameter; The elastic network regression model is trained based on the target regularization parameter to obtain the trained aging assessment model. Based on the trained aging assessment model, extracellular matrix aging-related proteins are identified, and an extracellular matrix aging assessment model is generated. Based on the extracellular matrix aging assessment model, an extracellular mechanism-predicted age is generated. The predictive performance of the extracellular matrix aging assessment model is evaluated using mean absolute error, and the model is validated by comparing the difference between the predicted age and the actual age.
7. The method according to claim 6, characterized in that, The formula for the objective function is as follows: ; in, Represented by the regression coefficient vector To optimize the object, find the model parameters that minimize the objective function; The output variable represents age; Represents a standardized protein expression matrix; Represents the regression coefficient vector; This represents the squared error term between the predicted age and the actual age. Indicates the penalty coefficient; This represents the ratio of the first regularization constraint; This represents the first norm of the regression coefficient vector; This represents the squared second norm of the regression coefficient vector.
8. The method according to claim 1, characterized in that, The step of correcting the difference between the predicted age and the actual age based on the extracellular matrix to obtain the extracellular matrix aging age difference includes: Based on the predicted age from the extracellular matrix and the actual age, sample pairing is performed to obtain age-paired data; Based on the age-paired data, a locally weighted scatter plot smoothing model is constructed to determine the baseline predicted age of the population corresponding to different actual ages. The initial aging age difference is obtained by calculating the difference between the extracellular matrix predicted age and the population baseline predicted age. The initial aging age difference is oriented and numerically corrected to obtain the extracellular matrix aging age difference.
9. The method according to claim 8, characterized in that, The calculation formula for the locally weighted scatter plot smoothing model is as follows: ; in, The locally weighted scatter plot smoothing method is represented in the first... Individual's actual age The fitted predicted age value obtained at this point is the baseline predicted age of the population; This indicates a local weighted scatter plot smoothing method; Indicates the first The actual age of each individual; Indicates the first The predicted age of an individual is output by the elastic network regression model; This indicates the proportion of neighboring data points used in local regression.
10. A plasma proteomics-based extracellular matrix aging assessment system, characterized in that, The system includes: The data acquisition module is used to acquire population sample data and plasma proteomics data to be tested. The population data processing module is used to preprocess the population sample data to obtain the target population dataset; The candidate protein extraction module is used to extract extracellular matrix candidate protein data from the plasma proteomics data based on a preset extracellular matrix protein list. The protein data processing module is used to perform quality control, missing value imputation, and expression level standardization on the extracellular matrix candidate protein data to obtain standardized protein expression data. The evaluation model construction module is used to construct an extracellular matrix aging evaluation model based on the target population dataset and the standardized protein expression data, obtain the predicted age of the extracellular matrix and evaluate the model performance. The age difference calculation module is used to correct the deviation between the predicted age and the actual age based on the extracellular matrix to obtain the extracellular matrix aging age difference. The result generation module is used to generate extracellular matrix aging assessment results based on the predicted age of the extracellular matrix, the aging age difference of the extracellular matrix, and model performance indicators.