A data processing method based on multi-omics features
Patent Information
- Application Number
- CN202611044861.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-18
AI Technical Summary
[0006]为了解决现有的技术问题,本发明提供了一种基于多组学特征的数据处理方法,通过整合核酸表达定量数据与蛋白质浓度多组学数据,经标准化消除不同指标量纲差异后,依托预设逻辑回归分类模型进行线性加权运算、概率转换及阈值判别,实现多组学数据的自动化分类判定,数据处理流程标准化、可批量自动化运算,运算效率高、成本可控,适配大规模样本多组学数据的批量分类筛查,弥补了单一组学数据信息维度不足、分类判别可靠性有限的缺陷,提升了多组学生物数据分类识别的整体性与精准度
获取待处理数据集;所述待处理数据集包括多个目标核酸表达定量数据和多个蛋白质浓度数据;对所述多个目标核酸表达定量数据和所述多个蛋白质浓度数据进行标准化处理,得到多个标准输入数据;将所述多个标准输入数据输入预设的逻辑回归分类模型中进行线性加权求和,计算得到判别值;将所述判别值转换为分类概率值;基于所述分类概率值和预设的概率阈值确定所述待处理数据集的数据类别。在本公开实施例中,通过整合核酸表达定量数据与蛋白质浓度多组学数据,经标准化消除不同指标量纲差异后,依托预设逻辑回归分类模型进行线性加权运算、概率转换及阈值判别,实现多组学数据的自动化分类判定,数据处理流程标准化、可批量自动化运算,运算效率高、成本可控,适配大规模样本多组学数据的批量分类筛查,弥补了单一组学数据信息维度不足、分类判别可靠性有限的缺陷,提升了多组学生物数据分类识别的整体性与精准度。
Smart Images

Figure CN122598748A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biological data processing technology, and in particular to a data processing method based on multi-omics features. Background Technology
[0002] Alzheimer's disease (AD) is a neurodegenerative disease with complex causes that leads to severe intellectual disability and is the most common type of dementia.
[0003] Over the past 20 years, significant progress has been made in the research of Alzheimer's disease (AD) biomarkers. Initially, the diagnosis of AD relied heavily on cerebrospinal fluid testing (such as Aβ42 and p-tau protein) and positron emission tomography (PET, such as Aβ-PET and Tau-PET). Although these methods have high diagnostic value, they are invasive examinations or expensive and have limited availability, which limits their application in large-scale screening.
[0004] In recent years, blood biomarkers have become a hot topic in the early diagnosis of Alzheimer's disease (AD) due to their minimally invasive, convenient, and large-scale application characteristics. For example, existing studies have found that protein biomarkers such as plasma p-tau181, p-tau217, and the Aβ42 / 40 ratio are strongly correlated with the core pathology of AD.
[0005] In the fields of bioinformatics, medical big data processing, and computer-aided analysis, collecting bodily fluid samples from examinees and extracting multi-source biological characteristic data from them, then using computer classification models to automatically classify and identify the data categories of the samples, is one of the core directions of current data science research. Existing classification algorithms often rely only on the data matrix of a single protein dimension. Due to the multi-network regulatory characteristics of complex biological systems, a single set of mathematical data cannot fully reflect the cascade of data characteristics of the examinee as a whole, resulting in machine learning models having severely insufficient specificity and sensitivity when faced with early subject data with unclear features. Summary of the Invention
[0006] To address existing technical problems, this invention provides a data processing method based on multi-omics features. By integrating quantitative nucleic acid expression data and protein concentration multi-omics data, and after standardization to eliminate differences in the dimensions of different indicators, linear weighting, probability transformation, and threshold discrimination are performed based on a preset logistic regression classification model. This achieves automated classification and determination of multi-omics data. The data processing flow is standardized, can be automated in batches, has high computational efficiency, and is cost-controllable. It is suitable for batch classification and screening of large-scale multi-omics data, making up for the shortcomings of insufficient information dimensions and limited reliability of classification and discrimination of single-omics data, and improving the overall integrity and accuracy of classification and identification of multi-omics biological data.
[0007] In a first aspect, embodiments of this application provide a data processing method based on multi-omics features, the method comprising: Obtain the dataset to be processed; the dataset to be processed includes quantitative data of expression of multiple target nucleic acids and concentration data of multiple proteins; The quantitative data of expression of the multiple target nucleic acids and the concentration data of the multiple proteins are standardized to obtain multiple standard input data. The multiple standard input data are input into a preset logistic regression classification model for linear weighted summation to calculate the discriminant value; The discriminant value is converted into a classification probability value; The data category of the dataset to be processed is determined based on the classification probability value and the preset probability threshold.
[0008] In an optional embodiment, the method further includes: Obtain a training sample set; the training sample set includes multiple sample nucleic acid data and corresponding data categories; Based on nucleic acid data from multiple samples, differential expression analysis was performed between different data categories to select multiple candidate nucleic acid data from the multiple sample nucleic acid data. Using the multiple candidate nucleic acid data as input features and the corresponding data categories as outputs, a random forest model is constructed for feature training to obtain feature importance scores for the multiple candidate nucleic acid data. Based on the feature importance score of the multiple candidate nucleic acid data, multiple target nucleic acid expression quantitative data are determined from the multiple candidate nucleic acid data.
[0009] In one optional embodiment, the quantitative data on the expression of the plurality of target nucleic acids include the expression levels of hsa-let-7d-5p, hsa-miR-20b-5p, hsa-miR-144-5p, hsa-miR-363-3p, and hsa-let-7e-5p.
[0010] In one optional embodiment, the differential expression analysis performed between different data categories based on multiple sample nucleic acid data, and the selection of multiple candidate nucleic acid data from the multiple sample nucleic acid data, includes: The t-test or rank-sum test was used to calculate the significant differences in nucleic acid data of the multiple samples between different data categories; If the significant difference value is greater than a preset significant threshold, the sample nucleic acid data is determined to be the candidate nucleic acid data.
[0011] In one optional embodiment, the standardization of the quantitative data of the expression of the plurality of target nucleic acids and the plurality of protein concentration data to obtain a plurality of standard input data includes: Based on the nucleic acid data of the multiple samples, a preset nucleic acid mean, a preset nucleic acid standard deviation, a preset protein mean, and a preset protein standard deviation are obtained; Based on the preset nucleic acid mean, the preset nucleic acid standard deviation, and the quantitative data of multiple target nucleic acid expression, a standard nucleic acid data set is obtained by standardizing the data. Based on the preset protein mean, the preset protein standard deviation, and the multiple protein concentration data, a standard protein concentration data is obtained by standardizing the data. The multiple standard input data are determined based on the multiple standard nucleic acid data and the multiple standard protein concentration data.
[0012] In an optional embodiment, determining the data category of the dataset to be processed based on the classification probability value and a preset probability threshold includes: If the classification probability value is greater than the probability threshold, the first type of data is determined to be the data category of the dataset to be processed. Alternatively, if the classification probability value is less than or equal to the probability threshold, the second type of data is determined to be the data category of the dataset to be processed.
[0013] In an optional embodiment, before inputting the plurality of standard input data into a preset logistic regression classification model for linear weighted summation to calculate the discriminant value, the method further includes: On the training sample set, a two-fold cross-validation combined with a hyperparameter grid search strategy is adopted, and the optimal classification performance evaluation index is used as the iteration termination condition for fitting training to determine the intercept and weight coefficients of the logistic regression classification model. The logistic regression classification model is determined based on the intercept and the weight coefficients.
[0014] Secondly, embodiments of this application provide a data processing apparatus based on multi-omics features, the apparatus comprising: The acquisition module is used to acquire the dataset to be processed; the dataset to be processed includes multiple target nucleic acid expression quantification data and multiple protein concentration data; The standardization module is used to standardize the quantitative data of the expression of the multiple target nucleic acids and the concentration data of the multiple proteins to obtain multiple standard input data. The calculation module is used to input the multiple standard input data into a preset logistic regression classification model for linear weighted summation and to calculate the discriminant value; The conversion module is used to convert the discrimination value into a classification probability value; The determination module is used to determine the data category of the dataset to be processed based on the classification probability value and a preset probability threshold.
[0015] Thirdly, embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the data processing method based on multi-omics features of the first aspect.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to implement the data processing method based on multi-omics features of the first aspect.
[0017] Fifthly, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the data processing method based on multi-omics features according to the first aspect.
[0018] The data processing method based on multi-omics features provided in this application has the following technical effects: A dataset to be processed is obtained; the dataset includes multiple target nucleic acid expression quantitative data and multiple protein concentration data; the multiple target nucleic acid expression quantitative data and the multiple protein concentration data are standardized to obtain multiple standard input data; the multiple standard input data are input into a preset logistic regression classification model for linear weighted summation to calculate the discriminant value; the discriminant value is converted into a classification probability value; the data category of the dataset to be processed is determined based on the classification probability value and a preset probability threshold. In this embodiment of the disclosure, by integrating nucleic acid expression quantitative data and protein concentration multi-omics data, after standardization to eliminate differences in the dimensions of different indicators, linear weighted calculation, probability transformation and threshold discrimination are performed based on a preset logistic regression classification model to achieve automated classification and determination of multi-omics data. The data processing flow is standardized, can be automated in batches, has high computational efficiency and controllable cost, and is suitable for batch classification and screening of large-scale sample multi-omics data. It makes up for the shortcomings of insufficient information dimensions and limited reliability of classification and discrimination of single omics data, and improves the overall integrity and accuracy of classification and identification of multi-omics biological data. Attached Figure Description
[0019] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a data processing method based on multi-omics features provided in an embodiment of this application. Figure 1 ; Figure 3 This is a flowchart illustrating a data processing method based on multi-omics features provided in an embodiment of this application. Figure 2 ; Figure 4 This is a flowchart illustrating a data processing method based on multi-omics features provided in an embodiment of this application. Figure 3 ; Figure 5 This is a schematic diagram of the structure of a data processing device based on multi-omics features provided in an embodiment of this application; Figure 6 This is a hardware structure block diagram of a server for a data processing method based on multi-omics features provided in an embodiment of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0023] Please see Figure 1 , Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application, including a data receiving device 101 and a server 102.
[0024] In one possible embodiment, the data receiving device 101 is a terminal acquisition device, mainly used to collect raw biological multi-omics data, specifically including raw nucleic acid sequencing data and protein biomarker concentration detection data, and to package and encapsulate the collected raw data and send it to the server 102; wherein, the data receiving device 101 can be adapted to various biological detection devices such as nucleic acid sequencers, flow cytometers, and chemiluminescence detectors, and can be compatible with different types of multi-omics raw data acquisition work.
[0025] In one possible embodiment, server 102 is configured with an independent data storage module, a data processing module, and a model calling module, and has integrated data processing capabilities such as data storage, data cleaning, feature filtering, model training, and data classification.
[0026] Specifically, server 102 acquires the dataset to be processed; the dataset includes multiple target nucleic acid expression quantification data and multiple protein concentration data; the multiple target nucleic acid expression quantification data and the multiple protein concentration data are standardized to obtain multiple standard input data; the multiple standard input data are input into a preset logistic regression classification model for linear weighted summation to calculate the discriminant value; the discriminant value is converted into a classification probability value; the data category of the dataset to be processed is determined based on the classification probability value and a preset probability threshold.
[0027] In this embodiment, by integrating nucleic acid expression quantitative data and protein concentration multi-omics data, and after standardization to eliminate differences in the dimensions of different indicators, linear weighting, probability transformation and threshold discrimination are performed based on a preset logistic regression classification model to achieve automated classification and judgment of multi-omics data. The data processing flow is standardized, can be automated in batches, has high computational efficiency and controllable cost, and is suitable for batch classification and screening of large-scale sample multi-omics data. It makes up for the shortcomings of insufficient information dimensions and limited reliability of classification and discrimination of single-omics data, and improves the overall integrity and accuracy of classification and recognition of multi-omics biological data.
[0028] The following describes a specific embodiment of a data processing method based on multi-omics features according to this application. Figure 2 This is a flowchart illustrating a data processing method based on multi-omics features provided in an embodiment of this application. Figure 1This specification provides method operation steps as shown in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual system or server products, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown in the embodiments or drawings... Figure 2 As shown, it may include: S201: Obtain the dataset to be processed; the dataset to be processed includes quantitative data of multiple target nucleic acid expression and multiple protein concentration data.
[0029] S202: Standardize the quantitative data of the expression of the multiple target nucleic acids and the concentration data of the multiple proteins to obtain multiple standard input data.
[0030] S203: Input the multiple standard input data into a preset logistic regression classification model and perform linear weighted summation to calculate the discriminant value.
[0031] S204: Convert the discriminant value into a classification probability value.
[0032] S205: Determine the data category of the dataset to be processed based on the classification probability value and the preset probability threshold.
[0033] Figure 3 This is a flowchart illustrating a data processing method based on multi-omics features provided in an embodiment of this application. Figure 2 The method may include: S301: Obtain the training sample set.
[0034] The training sample set includes multiple sample nucleic acid data and corresponding data categories.
[0035] In this embodiment, several subjects are pre-recruited and ex vivo plasma samples are collected as baseline samples. All baseline samples undergo uniform screening and preprocessing to ensure sample homogeneity. Nucleic acid sequencing and protein biomarker detection are performed on all samples, and corresponding sample nucleic acid data and protein concentration data are collected. Based on the sample data attributes, corresponding data categories are pre-defined, and all data are integrated to construct a training sample set. The training sample set is randomly divided into a validation subset and a discovery subset in a 3:7 ratio. The discovery subset is used for feature selection and model iterative training, while the validation subset is used to evaluate the model's classification performance.
[0036] S302: Based on multiple sample nucleic acid data, perform differential expression analysis between different data categories, and screen out multiple candidate nucleic acid data from the multiple sample nucleic acid data.
[0037] In this embodiment, the nucleic acid data is miRNA expression quantification data after being standardized by Counts Per Million (CPM). The original sequencing data is quantitatively statistically analyzed using the miRge3.0 tool. After obtaining the count value corresponding to each miRNA, a second standardization is performed to avoid the impact of sequencing depth differences on data accuracy.
[0038] Differential expression analysis is used to identify nucleic acid features that show significant numerical differences between different data categories, and high-value candidate features are screened. The specific steps are as follows: S3021: Use t-test or rank-sum test to calculate the significant difference values of the nucleic acid data of the multiple samples between different data categories.
[0039] In one optional embodiment, for nucleic acid data samples that conform to a normal distribution, an independent samples t-test is used to perform between-group difference analysis; for skewed data that do not conform to a normal distribution, a rank-sum test is used to perform difference analysis. By adapting the above two test methods to nucleic acid data of different distribution types, the p-value of each nucleic acid data item in the two types of sample data is accurately calculated, and the p-value is used as the significant difference value.
[0040] S3022: If the significance difference value is greater than the preset significance threshold, the sample nucleic acid data is determined to be the candidate nucleic acid data.
[0041] In this embodiment, the preset significance threshold can be customized according to the data screening accuracy requirements. Under normal working conditions, the threshold is set to 0.05. When the significance difference value corresponding to the nucleic acid data is greater than 0.05, it indicates that the nucleic acid data has significant differences in expression characteristics in different data categories and has classification and identification value, so it is included in the candidate nucleic acid data set; otherwise, the invalid nucleic acid feature is removed.
[0042] S303: Using the multiple candidate nucleic acid data as input features and the corresponding data categories as outputs, a random forest model is constructed for feature training to obtain feature importance scores for the multiple candidate nucleic acid data.
[0043] In this embodiment, a random forest feature selection model is built using the RandomForestClassifier module from the scikit-learn toolkit in Python. All candidate nucleic acid data obtained from the aforementioned selection are used as input features to the model, and pre-defined data categories are used as supervision labels. The models are then iteratively trained. After training, the feature importance score for each candidate nucleic acid data item is automatically calculated. A higher score indicates a greater influence of the nucleic acid feature on the data classification result, resulting in stronger classification and recognition capabilities.
[0044] S304: Based on the feature importance score of the multiple candidate nucleic acid data, determine multiple target nucleic acid expression quantitative data in the multiple candidate nucleic acid data.
[0045] In this embodiment, the SelectFromModel module in the scikit-learn toolkit is used to sort the candidate nucleic acid data in descending order based on feature importance scores, and the top five nucleic acid data are selected as the final target nucleic acid expression quantification data.
[0046] In one optional embodiment, the quantitative data on the expression of multiple target nucleic acids specifically include the CPM-normalized expression levels of hsa-let-7d-5p, hsa-miR-20b-5p, hsa-miR-144-5p, hsa-miR-363-3p, and hsa-let-7e-5p; simultaneously, two types of protein feature data are matched, namely p-tau181 protein concentration and Aβ42 / Aβ40 ratio, wherein the Aβ42 / Aβ40 ratio is calculated from the measured Aβ42 and Aβ40 protein concentrations, and finally, the seven features are integrated as input features for model building.
[0047] S305: On the training sample set, a two-fold cross-validation combined with a hyperparameter grid search strategy is adopted, and the optimal classification performance evaluation index is used as the termination condition for iteration to perform fitting training, thereby determining the intercept and weight coefficients of the logistic regression classification model.
[0048] In this embodiment, the LogisticRegression module in the scikit-learn toolkit is called to initialize the basic logistic regression model, and the GridSearchCV grid search module is used to complete the global optimization of hyperparameters. The training process adopts a two-fold cross-validation mode, which divides the training subset equally into two validation units, and performs training and validation in turn to reduce the risk of overfitting caused by training on a single dataset.
[0049] This model uses the F1 score as the core optimization objective, while simultaneously considering accuracy, precision, sensitivity, specificity, and the area under the receiver operating characteristic curve (AUC). It continuously iterates and adjusts the model's hyperparameters, feature weight coefficients, and intercept parameters until the F1 score reaches its peak, at which point the model training is terminated.
[0050] S306: Determine the logistic regression classification model based on the intercept and the weight coefficients.
[0051] In this embodiment of the application, after iterative training is completed, the intercept and weight coefficients of each feature in the optimal state are locked to generate a fixed version of the logistic regression classification model.
[0052] The model calculation formula is as follows: Logit=1.9596+0.6831×zscore1+0.7749×zscore2+1.4616×zscore3+0.8212×zscore4+0.3472×zscore5+2.5722×zscore6+0.1929×zscore7.
[0053] In the formula, zscore1-zscore7 correspond to the standardized data of the five target nucleic acid data, p-tau181 protein concentration, and Aβ42 / Aβ40 ratio, respectively.
[0054] After validation subset testing, the model achieved an overall classification accuracy of 0.8333 and an AUC of 0.8536, demonstrating excellent multi-omics data classification capabilities.
[0055] Figure 4 This is a flowchart illustrating a data processing method based on multi-omics features provided in an embodiment of this application. Figure 3 The method may include: S401: Obtain the dataset to be processed.
[0056] The dataset to be processed includes quantitative data on the expression of multiple target nucleic acids and concentration data on multiple proteins.
[0057] In this embodiment of the application, the raw detection data of the sample to be processed is collected by a biological detection device. The raw data includes the raw data of target nucleic acid expression and the raw data of the concentrations of three types of proteins: p-tau181, Aβ42, and Aβ40.
[0058] The raw nucleic acid data were standardized using CPM to obtain quantitative data on target nucleic acid expression. The Aβ42 / Aβ40 ratio was calculated based on the Aβ42 and Aβ40 data. All data were then integrated to construct a dataset for processing. This dataset consists of five sets of quantitative target nucleic acid expression data, one set of protein concentration data, and one set of protein ratio data. The nucleic acid data can be obtained through sequencing, quantitative polymerase chain reaction (PCR), digital PCR, microarray detection, etc., while the protein concentration data can be obtained through flow cytometry, chemiluminescence, enzyme-linked immunosorbent assay (ELISA), etc.
[0059] S402: Standardize the quantitative data of the expression of the multiple target nucleic acids and the concentration data of the multiple proteins to obtain multiple standard input data.
[0060] In this embodiment, because the numerical ranges and dimensional properties of nucleic acid expression quantification data, protein concentration data, and protein ratio data differ significantly, directly inputting them into the model would cause weight shifts and affect classification accuracy. Therefore, the Z-score standardization algorithm is used to unify the data dimensions. The standardization formula is: zscore=(x-mean) / std; where x is the original feature detection value, mean is the mean of the corresponding feature training set, and std is the standard deviation of the corresponding feature training set. The specific steps are as follows: S4021: Based on the nucleic acid data of the multiple samples, obtain the preset nucleic acid mean, preset nucleic acid standard deviation, preset protein mean, and preset protein standard deviation.
[0061] In one optional embodiment, the mean and standard deviation are both obtained from the raw training data during the model training phase, and the fixed parameters for each feature are as follows: hsa-let-7d-5p: mean=2532.558577, std=1014.729857; hsa-miR-20b-5p: mean=208.981396, std=113.686776; hsa-miR-144-5p: mean=82.676848, std=58.088469; hsa-miR-363-3p: mean=83.173828, std=54.982848; hsa-let-7e-5p: mean=1272.415786, std=1395.684122; p-tau181: mean=1.245266, std=0.851899; Aβ42 / Aβ40: mean=0.119052, std=0.058771.
[0062] S4022: Based on the preset nucleic acid mean, the preset nucleic acid standard deviation, and the quantitative data of multiple target nucleic acid expression, standardization processing is performed to obtain multiple standard nucleic acid data.
[0063] In this embodiment, the quantitative data of nucleic acid expression of the five targets are substituted into the Z-score formula, and the corresponding nucleic acid mean and standard deviation are matched to calculate the standardized nucleic acid data, namely zscore1 to zscore5.
[0064] S4023: Based on the preset protein mean, the preset protein standard deviation, and the multiple protein concentration data, perform standardization processing to obtain multiple standard protein concentration data.
[0065] In this embodiment, the p-tau181 concentration data and Aβ42 / Aβ40 ratio data are standardized, and the corresponding protein mean and standard deviation are matched to obtain the standard protein concentration data zscore6 and the standard protein ratio data zscore7.
[0066] S4024: Determine the multiple standard input data based on the multiple standard nucleic acid data and the multiple standard protein concentration data.
[0067] In this embodiment of the application, the calculated zscore1, zscore2, zscore3, zscore4, zscore5, zscore6, and zscore7 are systematically integrated to form 7-dimensional standard input data with uniform dimensions and consistent units, which are used as input parameters for the logistic regression model.
[0068] S403: Input the multiple standard input data into a preset logistic regression classification model and perform linear weighted summation to calculate the discriminant value.
[0069] In this embodiment, 7-dimensional standard input data is imported into the solidified logistic regression classification model. The model calls the built-in weight coefficients and intercept parameters to complete the linear weighted summation operation and output the corresponding Logit discriminant value. This discriminant value is a dimensionless intermediate operation parameter used for subsequent probability transformation.
[0070] S404: Convert the discriminant value into a classification probability value.
[0071] In this embodiment, the sigmoid activation function is used to map and transform the Logit discriminant value, converting the discriminant value without a fixed interval into a classification probability value in the range of 0 to 1. The magnitude of the probability value represents the confidence level of the current dataset belonging to the corresponding category.
[0072] S405: Determine the data category of the dataset to be processed based on the classification probability value and the preset probability threshold.
[0073] In this embodiment, staff can customize a probability threshold according to the data classification accuracy requirements. By comparing the probability value with the threshold, the data category classification is completed. The specific determination logic is as follows: S4051: If the classification probability value is greater than the probability threshold, determine the first type of data as the data category of the dataset to be processed.
[0074] S4052: Alternatively, if the classification probability value is less than or equal to the probability threshold, determine the second type of data as the data category of the dataset to be processed.
[0075] In this embodiment, the data processing method completes model training and data classification based on multi-omics fusion features, which differs from traditional single-omics data processing schemes. It can mine the intrinsic features of data from the dual dimensions of nucleic acid regulation and protein expression. The entire method does not require large imaging equipment or invasive sampling operations. Data collection can be completed using only conventional biological detection methods. Combined with automated data processing workflows, it significantly reduces data processing costs. At the same time, the model has undergone cross-validation and hyperparameter optimization, resulting in strong classification stability. It can efficiently complete the classification and screening of massive multi-omics biological data and is suitable for various application scenarios such as biological research and sample data screening.
[0076] This application also provides a data processing apparatus based on multi-omics features. Figure 5 This is a schematic diagram of the structure of a data processing device based on multi-omics features provided in an embodiment of this application, as shown below. Figure 5 As shown, the device 500 includes: The acquisition module 510 is used to acquire the dataset to be processed; the dataset to be processed includes multiple target nucleic acid expression quantification data and multiple protein concentration data; The standardization module 520 is used to standardize the quantitative data of the expression of the multiple target nucleic acids and the multiple protein concentration data to obtain multiple standard input data. The calculation module 530 is used to input the multiple standard input data into a preset logistic regression classification model for linear weighted summation and to calculate the discriminant value; The conversion module 540 is used to convert the discrimination value into a classification probability value; The determination module 550 is used to determine the data category of the dataset to be processed based on the classification probability value and a preset probability threshold.
[0077] In one alternative implementation, it further includes: The first acquisition module is used to acquire a training sample set; the training sample set includes multiple sample nucleic acid data and corresponding data categories. The screening module is used to perform differential expression analysis between different data categories based on multiple sample nucleic acid data, and to screen out multiple candidate nucleic acid data from the multiple sample nucleic acid data. The training module is used to construct a random forest model for feature training by taking the multiple candidate nucleic acid data as input features and the corresponding data categories as output, and to obtain feature importance scores for the multiple candidate nucleic acid data. The first determining module is used to determine multiple target nucleic acid expression quantification data from the multiple candidate nucleic acid data based on the feature importance score of the multiple candidate nucleic acid data.
[0078] In one optional embodiment, the quantitative data on the expression of the plurality of target nucleic acids include the expression levels of hsa-let-7d-5p, hsa-miR-20b-5p, hsa-miR-144-5p, hsa-miR-363-3p, and hsa-let-7e-5p.
[0079] In an optional embodiment, it further includes: The first calculation module is used to calculate the significant difference values of the nucleic acid data of the multiple samples between different data categories using t-test or rank-sum test; The second determining module is used to determine the sample nucleic acid data as the candidate nucleic acid data if the significant difference value is greater than a preset significant threshold.
[0080] In an optional embodiment, it further includes: The second calculation module is used to obtain a preset nucleic acid mean, a preset nucleic acid standard deviation, a preset protein mean, and a preset protein standard deviation based on the multiple sample nucleic acid data; The first standardization module is used to perform standardization processing based on the preset nucleic acid mean, the preset nucleic acid standard deviation and the multiple target nucleic acid expression quantitative data to obtain multiple standard nucleic acid data. The second standardization module is used to perform standardization processing based on the preset protein mean, the preset protein standard deviation and the multiple protein concentration data to obtain multiple standard protein concentration data. The third determining module is used to determine the multiple standard input data based on the multiple standard nucleic acid data and the multiple standard protein concentration data.
[0081] In an optional embodiment, it further includes: The fourth determining module is used to determine the first type of data as the data category of the dataset to be processed if the classification probability value is greater than the probability threshold. The fifth determining module is used to determine the second type of data as the data category of the dataset to be processed, either if the classification probability value is less than or equal to the probability threshold.
[0082] In an optional embodiment, it further includes: The sixth determination module is used to perform fitting training on the training sample set using a two-fold cross-validation combined with a hyperparameter grid search strategy, with the optimal classification performance evaluation index as the iteration termination condition, to determine the intercept and weight coefficients of the logistic regression classification model. The seventh determining module is used to determine the logistic regression classification model based on the intercept and the weight coefficients.
[0083] The apparatus and method embodiments in this application are based on the same application concept.
[0084] The methods and embodiments provided in this application can be executed on a computer terminal, server, or similar computing device. Taking running on a server as an example, Figure 6 This is a hardware structure block diagram of a server for a data processing method based on multi-omics features provided in an embodiment of this application. Figure 6 As shown, the server 600 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 610 (CPUs 610 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 630 for storing data, and one or more storage media 620 (e.g., one or more mass storage devices) for storing application programs 623 or data 622. The memory 630 and storage media 620 may be temporary or persistent storage. The program stored in the storage media 620 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 610 may be configured to communicate with the storage media 620 and execute the series of instruction operations stored in the storage media 620 on the server 600. Server 600 may also include one or more power supplies 660, one or more wired or wireless network interfaces 650, one or more input / output interfaces 640, and / or one or more operating systems 621, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0085] The input / output interface 640 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 600. In one example, input / output interface 640 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, input / output interface 640 may be a radio frequency (RF) module used for wireless communication with the Internet.
[0086] Those skilled in the art will understand that Figure 6The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 600 may also include... Figure 6 The more or fewer components shown, or having the same Figure 6 The different configurations shown.
[0087] This application provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The processor loads and executes the at least one instruction, at least one program, code set, or instruction set to implement the above-described data processing method.
[0088] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in a server to store at least one instruction, at least one program, code set, or instruction set related to implementing a data processing method based on multi-omics features in the method embodiment. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the aforementioned data processing method based on multi-omics features.
[0089] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0090] As can be seen from the embodiments of the data processing method, apparatus, electronic device or storage medium based on multi-omics features provided in this application, the present application obtains a dataset to be processed; the dataset to be processed includes multiple target nucleic acid expression quantitative data and multiple protein concentration data; the multiple target nucleic acid expression quantitative data and the multiple protein concentration data are standardized to obtain multiple standard input data; the multiple standard input data are input into a preset logistic regression classification model for linear weighted summation to calculate a discriminant value; the discriminant value is converted into a classification probability value; and the data category of the dataset to be processed is determined based on the classification probability value and a preset probability threshold. In this embodiment, by integrating nucleic acid expression quantitative data and protein concentration multi-omics data, and after standardization to eliminate differences in the dimensions of different indicators, linear weighting, probability transformation and threshold discrimination are performed based on a preset logistic regression classification model to achieve automated classification and judgment of multi-omics data. The data processing flow is standardized, can be automated in batches, has high computational efficiency and controllable cost, and is suitable for batch classification and screening of large-scale sample multi-omics data. It makes up for the shortcomings of insufficient information dimensions and limited reliability of classification and discrimination of single-omics data, and improves the overall integrity and accuracy of classification and recognition of multi-omics biological data.
[0091] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0092] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0093] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0094] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A data processing method based on multi-omics features, characterized in that, include: Obtain the dataset to be processed; the dataset to be processed includes quantitative data of expression of multiple target nucleic acids and concentration data of multiple proteins; The quantitative data of expression of the multiple target nucleic acids and the concentration data of the multiple proteins are standardized to obtain multiple standard input data. The multiple standard input data are input into a preset logistic regression classification model for linear weighted summation to calculate the discriminant value; The discriminant value is converted into a classification probability value; The data category of the dataset to be processed is determined based on the classification probability value and the preset probability threshold. 2.The data processing method based on multi-omics characteristics of claim 1, wherein, The method further includes: Obtain a training sample set; the training sample set includes multiple sample nucleic acid data and corresponding data categories; Based on nucleic acid data from multiple samples, differential expression analysis was performed between different data categories to select multiple candidate nucleic acid data from the multiple sample nucleic acid data. Using the multiple candidate nucleic acid data as input features and the corresponding data categories as outputs, a random forest model is constructed for feature training to obtain feature importance scores for the multiple candidate nucleic acid data. Based on the feature importance score of the multiple candidate nucleic acid data, multiple target nucleic acid expression quantitative data are determined from the multiple candidate nucleic acid data.
3. The data processing method based on multi-omics features according to claim 2, characterized in that, The quantitative data on the expression of multiple target nucleic acids include the expression levels of hsa-let-7d-5p, hsa-miR-20b-5p, hsa-miR-144-5p, hsa-miR-363-3p, and hsa-let-7e-5p.
4. The data processing method based on multi-omics features according to claim 2, characterized in that, The method involves differential expression analysis across different data categories based on multiple sample nucleic acid data, and screening out multiple candidate nucleic acid data from the multiple sample nucleic acid data, including: The t-test or rank-sum test was used to calculate the significant differences in nucleic acid data of the multiple samples between different data categories; If the significant difference value is greater than a preset significant threshold, the sample nucleic acid data is determined to be the candidate nucleic acid data.
5. The data processing method based on multi-omics features according to claim 2, characterized in that, The standardization process is performed on the quantitative data of the expression of the multiple target nucleic acids and the multiple protein concentration data to obtain multiple standard input data, including: Based on the nucleic acid data of the multiple samples, a preset nucleic acid mean, a preset nucleic acid standard deviation, a preset protein mean, and a preset protein standard deviation are obtained; Based on the preset nucleic acid mean, the preset nucleic acid standard deviation, and the quantitative data of multiple target nucleic acid expression, a standard nucleic acid data set is obtained by standardizing the data. Based on the preset protein mean, the preset protein standard deviation, and the multiple protein concentration data, a standard protein concentration data is obtained by standardizing the data. The multiple standard input data are determined based on the multiple standard nucleic acid data and the multiple standard protein concentration data.
6. The data processing method based on multi-omics features according to claim 1, characterized in that, Determining the data category of the dataset to be processed based on the classification probability value and a preset probability threshold includes: If the classification probability value is greater than the probability threshold, the first type of data is determined to be the data category of the dataset to be processed. Alternatively, if the classification probability value is less than or equal to the probability threshold, the second type of data is determined as the data category of the dataset to be processed.
7. The data processing method based on multi-omics features according to claim 2, characterized in that, Before inputting the multiple standard input data into a preset logistic regression classification model for linear weighted summation to calculate the discriminant value, the following steps are also included: On the training sample set, a two-fold cross-validation combined with a hyperparameter grid search strategy is adopted, and the optimal classification performance evaluation index is used as the iteration termination condition for fitting training to determine the intercept and weight coefficients of the logistic regression classification model. The logistic regression classification model is determined based on the intercept and the weight coefficients.
8. A data processing device based on multi-omics features, characterized in that, The device includes: The acquisition module is used to acquire the dataset to be processed; the dataset to be processed includes multiple target nucleic acid expression quantification data and multiple protein concentration data; The standardization module is used to standardize the quantitative data of the expression of the multiple target nucleic acids and the concentration data of the multiple proteins to obtain multiple standard input data. The calculation module is used to input the multiple standard input data into a preset logistic regression classification model for linear weighted summation and to calculate the discriminant value; The conversion module is used to convert the discrimination value into a classification probability value; The determination module is used to determine the data category of the dataset to be processed based on the classification probability value and a preset probability threshold.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the data processing method based on multi-omics features as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the data processing method based on multi-omics features as described in any one of claims 1-7.