Treatment response data processing system based on multiple protein detection technology

Through the data processing system of multiple protein detection technology, the drug treatment-sensitive subgroups are identified and key protein markers are screened out, solving the problems of subgroup identification and characteristic variable selection in the prior art, and improving the accuracy of treatment response research.

CN120356702APending Publication Date: 2025-07-22TIANJIN CITY THIRD CENT HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510416421.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to identify subgroups of patients who are sensitive to drug treatment, resulting in challenges in precise treatment, and existing omics data processing methods are not suitable for treatment response studies.

Method used

A therapeutic response data processing system based on multiple protein detection technology, including data preprocessing, characteristic variable selection, efficacy quantification, subgroup division and subgroup evaluation modules, is used to identify sensitive subgroups and screen protein markers that reflect the therapeutic effect through singular value decomposition and principal component analysis.

Benefits of technology

The precise identification and screening of the patient subgroup and the protein markers that best reflect the therapeutic effect were achieved, improving the accuracy of treatment response research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356702A_ABST
    Figure CN120356702A_ABST
Patent Text Reader

Abstract

The invention provides a treatment response data processing system based on a multiple protein detection technology, and the system comprises a data preprocessing module which is used for preprocessing original data; and the characteristic variable selection module is used for screening the protein indexes to obtain the protein marker with the characteristic variable representation curative effect after screening. The curative effect quantification module is used for acquiring a change absolute value of the first main component as a curative effect quantification index; and the subgroup division module is used for acquiring subgroups sensitive to treatment in the patient. And the subgroup evaluation module is used for evaluating whether the subgroups acquired in the subgroup division module are reasonable or not. The method has the beneficial effects that the treatment response is evaluated through multiple protein detection data, and two key problems related to a multiple protein detection technology in the prior art are solved. A subgroup sensitive to treatment in a patient can be identified, namely, the subgroup is identified, and a protein marker which can best reflect the treatment effect can be screened, namely, the characteristic variable is selected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and more particularly to a treatment response data processing system based on a multiplex protein detection technology. Background Art

[0002] Patient blood samples collected in clinical trials often need to analyze multiple proteins, such as inflammatory factors, tumor markers, etc. The multiplex protein detection technology can analyze hundreds or even thousands of target proteins simultaneously from a small amount of samples. Taking the Proximity Extension Assay (PEA) technology as an example, it combines antibody-based immunoassay with polymerase chain reaction, and can perform high-throughput quantitative detection of low-abundance functional proteins, cytokines or biomarkers in blood, and has the characteristics of high sensitivity, wide linear range, good repeatability, etc.

[0003] The Swedish company Olink Proteomics AB has commercialized this technology and developed a series of targeted protein panels for various scenarios such as basic research, drug development and clinical trials. For example, its "Olink Target 96" series of products can perform high-throughput detection of 80-90 proteins related to specific diseases such as tumors, cardiovascular diseases, etc. or physiological processes such as inflammatory responses, immune responses, etc. in samples. Each run can simultaneously detect 96 samples, including 88 test samples and 8 quality control samples.

[0004] Based on the multiplex protein detection technology can be used in a variety of clinical fields, one of which is to evaluate the response of patients to drug treatment. For example, psoriasis patients receive biological agent treatment. Before the clinical symptoms are significantly improved, by detecting inflammatory factors or immune response indicators in the blood, the response of the body to the drug can be obtained more sensitively. For the design of treatment response research protocols, although patients are also divided into two groups before and after treatment, in the processing of experimental data, it is not possible to simply find differential proteins by comparing the two groups, because some patients may be insensitive to the treatment, and there is no difference in their pre- and post-treatment results; this is a characteristic that treatment response research is different from general case-control, such as the comparison study between disease and normal population, and most of the current statistical analysis methods based on omics data are only applicable to the latter.

[0005] Since patients may have heterogeneity in the efficacy of the same drug or treatment regimen, identifying sensitive subgroups is one of the challenges faced by precision medicine in clinical practice. Summary of the Invention

[0006] The present invention overcomes the deficiencies in the prior art and provides a treatment response data processing system based on multiplex protein detection technology.

[0007] The object of the present invention is achieved by the following technical solutions.

[0008] A treatment response data processing system based on multiplex protein detection technology includes:

[0009] A data preprocessing module for preprocessing the original data;

[0010] A feature variable selection module for calculating the p-values of the changes of each protein in the two groups before and after treatment, arranging the data according to the p-values, gradually retaining the data of the proteins with the smallest p-values in the data according to the set retention quantity or percentage value, performing principal component analysis on the data after each reduction of variables, obtaining the principal component scores of each sample, recalculating the p-values of the changes of the retained protein data group in the two groups before and after treatment and sorting the data in ascending order; repeating the above steps until the number of variable data is reduced to the retention percentage value of the original data quantity, and taking the variable data contained in the data at this time as the screened feature variables and as the protein markers for characterizing the curative effect;

[0011] A curative effect quantification module for performing principal component analysis on the data after variable selection, calculating the change value data of the first principal component before and after treatment of each patient sample, and taking the change value of the first principal component as the curative effect quantification index;

[0012] A subgroup division module for arranging the absolute values of the changes of the first principal component of each patient in descending order of scores, first taking the top 3 patients with the highest scores as a subgroup, calculating the number of differentially expressed proteins before and after treatment of the patients in this subgroup, then adding 1 patient in turn, recalculating the number of differentially expressed proteins before and after treatment of the new subgroup of patients, and repeating this step until when the number of differentially expressed proteins before and after treatment of the patients in the subgroup reaches the highest value, defining the subgroup at this time as the main response group;

[0013] A subgroup evaluation module for evaluating whether the subgroup obtained in the subgroup division module is reasonable.

[0014] The working steps of the data preprocessing module are as follows:

[0015] Dividing the original data K(ij) by the average value of its row to obtain a ratio value matrix R(ij),

[0016] R(ij) = K(ij) / [(∑K(i)) / M]

[0017] In the ratio value matrix R(ij), taking the median of each column value R(j) as the correction coefficient S(j) between samples,

[0018] S(j) = Median[R(j)]

[0019] The original data K(ij) is divided by the coefficient column by column to obtain the corrected data K'(ij).

[0020] K'(ij) = K(ij) / S(j)

[0021] Among them, the original data is a measurement value matrix K(ij) with N rows and M columns, where i = 1,...N represents proteins and j = 1,...M represents samples. The samples include two groups of data before and after the treatment of each patient; K(i) is the row data and K(j) is the column data.

[0022] In the process of gradually retaining the p-value data in the feature variable selection module, the first 100%, 50%, 25%, 10%, and 5% of the data are selected and retained in sequence.

[0023] The percentage of the number of variable data in the feature variable selection module reduced to the original data number is 5%.

[0024] The method steps of principal component analysis in the feature variable selection module are as follows:

[0025] Through singular value decomposition, the processed data X is decomposed into three matrices.

[0026]

[0027] Among them, U is the left singular vector matrix; S is the singular value diagonal matrix; is the transpose of the right singular vector matrix; the data X is the corrected data K'(ij);

[0028] Calculate the principal component scores. The principal component score calculation formula is

[0029] Score = U(n - 1) 1 / 2

[0030] Among them, n is the number of columns of U, and the first column Score i is the i-th principal component score.

[0031] The calculation formula for the change value of the first principal component is:

[0032] Ds = |Score 1 [B] – Score 1 [A]|

[0033] Among them, Ds is the change value of the first principal component (Delta score, Ds); B is before treatment and A is after treatment.

[0034] The working steps of the subgroup evaluation module are as follows:

[0035] Set the total number of patients as N and the number of patients in the subgroup as M;

[0036] Sample the preprocessed data multiple times, with the number of patients drawn each time being M, to form random subgroups;

[0037] Calculate the number of differentially expressed proteins before and after treatment within each subgroup drawn each time;

[0038] Judge whether the number of differentially expressed proteins in the subgroup obtained through discrimination is greater than the number of differentially expressed proteins in 95% of the random subgroups. When the conditions are met, it is determined that the obtained subgroup is reasonable.

[0039] Calculate the number of differentially expressed proteins before and after treatment within each subgroup drawn each time by means of paired t-test.

[0040] The beneficial effects of the present invention are as follows: This solution evaluates treatment response through multiplex protein detection data, and solves two key problems related to multiplex protein detection technology in the prior art. One is to identify the subgroup of patients who are more sensitive to treatment, that is, subgroup identification; the other is to screen out the protein markers that can best reflect the treatment effect, that is, variable selection. Description of the Drawings

[0041] Figure 1 is the working step diagram of the module of the present invention;

[0042] Figure 2 is the processed result data of the data preprocessing module in this embodiment;

[0043] Figure 3 is the processed result data of the variable selection module in this embodiment;

[0044] Figure 4 is the processed result data of the efficacy quantification module in this embodiment;

[0045] Figure 5 is the relationship diagram between the division of the main and secondary response groups and the number of differentially expressed proteins in this example;

[0046] Figure 6 is the processed result data of the subgroup evaluation module in this embodiment. Detailed Embodiment

[0047] The technical solution of the present invention will be further described below through specific embodiments.

[0048] Embodiment

[0049] In this embodiment, the data samples are patients with moderate to severe psoriasis who have been clinically evaluated by a doctor and received treatment with secukinumab injection; a total of 22 patients (P01-P22) were enrolled, and blood samples were collected before treatment and 4 weeks after treatment. The changes of 92 inflammatory factors in the patients' blood were analyzed using PEA-based multiplex protein detection technology; the purpose is to screen protein markers related to treatment response, construct efficacy evaluation indicators, and identify subgroups of the population that benefited more.

[0050] This system is used to process sample data.

[0051] The working steps of the data preprocessing module are:

[0052] Divide the original data K(ij) by the average value of its row to obtain the ratio value matrix R(ij),

[0053] R(ij)=K(ij) / [(∑K(i)) / M]

[0054] In the ratio value matrix R(ij), take the median of each column value R(j) as the correction coefficient S(j) between samples.

[0055] S(j)=Median[R(j)]

[0056] Divide the original data K(ij) by the coefficients column by column to get the corrected data K'(ij).

[0057] K'(ij)=K(ij) / S(j)

[0058] The original data is a measurement matrix K(ij) with N rows and M columns, where i=1,...N represents protein, j=1,...M represents sample, and the sample includes two sets of data for each patient before and after treatment; K(i) is row data and K(j) is column data.

[0059] During the sample collection, transportation and pre-processing stages, various uncontrollable factors may cause slight degradation of proteins in the samples to varying degrees; therefore, the data preprocessing module is used to preprocess the data and correct the median protein expression levels between samples to a similar level to minimize the degradation that may occur during the sample collection stage. Figure 2 As shown in FIG. 1 , the figure is the processed sample data in this embodiment.

[0060] A feature variable selection module, which is used to calculate the p-values of the changes of each protein in the two groups before and after treatment, arrange the data according to the p-values, gradually retain the data of the proteins with smaller p-values in the data, perform principal component analysis on the data after each variable reduction, obtain the principal component scores of each sample, recalculate the p-values of the changes of the retained protein data group in the two groups before and after treatment, and sort the data in ascending order; repeat the above steps until the number of variable data is reduced to the retention percentage value of the original data quantity, take the variables included in the data at this time as the screened feature variables, and take the first principal component as the protein biomarker representing the curative effect.

[0061] The method steps of the principal component analysis in the feature variable selection module are as follows:

[0062] Through singular value decomposition, decompose the processed data X into three matrices

[0063]

[0064] where U is the left singular vector matrix; S is the singular value diagonal matrix; is the transpose of the right singular vector matrix;

[0065] Calculate the principal component scores. The principal component score calculation formula is

[0066] Score = U(n - 1) 1 / 2

[0067] where n is the number of columns of U, and the first column Score i is the i-th principal component score.

[0068] Furthermore, in the process of gradually retaining the data of the proteins with smaller p-values in the feature variable selection module, the first 100%, 50%, 25%, 10% and 5% of the data are sequentially selected for retention. The retention percentage value when the number of variable data in the feature variable selection module is reduced to the original data quantity is 5%.

[0069] As Figure 3 shown, this figure is the processing result of the sample data in this embodiment.

[0070] An efficacy quantification module, which is used to perform principal component analysis on the data after variable selection, calculate the first principal component change value data of each patient sample before and after treatment, take the first principal component change value data as the efficacy quantification index, and the larger the absolute value of the change of the first principal component value, the more obvious the response to treatment.

[0071] The calculation formula for the change of the first principal component value is:

[0072] Ds = |Score 1[B]–Score 1 [A]|

[0073] Among them, Ds is the change value of the first principal component (Delta score, Ds); B is before treatment, and A is after treatment.

[0074] In this embodiment, the processing result of the efficacy quantification module for the sample data is as Figure 4 shown.

[0075] The subgroup division module is used to arrange the absolute values of the changes in the first principal component of each patient in descending order of scores. First, the top 3 patients in the ranking are used as a subgroup, and the number of differentially expressed proteins before and after treatment of the patients in this subgroup is calculated. Then, one patient is added in turn, and the number of differentially expressed proteins before and after treatment of the new subgroup of patients is recalculated. This step is repeated until the number of differentially expressed proteins before and after treatment of the patients in the subgroup reaches the highest value. At this time, the subgroup is defined as the main response group.

[0076] In this embodiment, when the number of patients included in the subgroup reaches 9, the number of differentially expressed proteins before and after treatment of the patients in the subgroup reaches the highest value of 31. At this time, it is determined that this subgroup is the main response subgroup with a large response to treatment, as Figure 5 shown.

[0077] The subgroup evaluation module is used to evaluate whether the subgroup obtained in the subgroup division module is reasonable.

[0078] The working steps of the subgroup evaluation module are as follows:

[0079] Set the total number of patients as N and the number of subgroup patients as M;

[0080] Sample the preprocessed data multiple times, and the number of patients drawn each time is M, forming a random subgroup;

[0081] Calculate the number of differentially expressed proteins before and after treatment within each subgroup drawn each time;

[0082] Judge whether the number of differentially expressed proteins in the subgroup obtained by discrimination is greater than the number of differentially expressed proteins in 95% of the random subgroups. When the condition is met, it is judged that the obtained subgroup is reasonable.

[0083] In this embodiment, a total of 22 patients were enrolled in the sample study, and 9 people were randomly sampled from them, with a total of 497,420 possible combinations.

[0084] To prove that the subgroup obtained by discrimination is different from the subgroup obtained by random sampling, 1000 simulations were performed with 9 samples each time, and the number of differentially expressed proteins before and after treatment within each random subgroup was calculated. The subgroup obtained by discrimination contained 31 differentially expressed proteins, while among the 1000 randomly sampled subgroups, only 4 had more than 31 differentially expressed proteins, that is, the p-value was 0.004. Therefore, it can be considered that the main response subgroup obtained by discrimination is statistically significant, and the module operation results are as Figure 6 shown.

[0085] The above has described the embodiments of the present invention in detail, but the content described is only the preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A treatment response data processing system based on multiplex protein detection technology, characterized in that, Including: A data preprocessing module for preprocessing the original data; A feature variable selection module for calculating the p-values of the changes of each protein in the two groups before and after treatment, arranging the data according to the p-values, gradually retaining the data of the proteins with the smallest p-values in the data according to the set retention quantity or percentage value, performing principal component analysis on the data after each variable reduction, obtaining the principal component scores of each sample, recalculating the p-values of the changes of the retained protein data group in the two groups before and after treatment and sorting the data in ascending order; repeating the above steps until the number of variable data is reduced to the retention percentage value of the original data quantity, taking the variable data included in the data at this time as the screened feature variables and as the protein markers characterizing the curative effect; A curative effect quantification module for performing principal component analysis on the data after variable selection, calculating the change value data of the first principal component before and after treatment of each patient sample, and taking the change value of the first principal component as the curative effect quantification index; A subgroup division module for arranging the absolute values of the changes of the first principal component of each patient in descending order of scores. First, taking the top 3 patients with the highest scores as a subgroup, calculating the number of differentially expressed proteins before and after treatment of the patients in this subgroup, and then adding 1 patient in turn, recalculating the number of differentially expressed proteins before and after treatment of the new subgroup of patients, repeating this step until when the number of differentially expressed proteins before and after treatment of the patients in the subgroup reaches the highest value, defining the subgroup at this time as the main response group; A subgroup evaluation module for evaluating whether the subgroup obtained in the subgroup division module is reasonable.

2. The treatment response data processing system based on the multiplex protein detection technology according to claim 1, wherein The working steps of the data preprocessing module are as follows: Dividing the original data K(ij) by the average value of its row to obtain a ratio value matrix R(ij), R(ij) = K(ij) / [(∑K(i)) / M] In the ratio value matrix R(ij), taking the median of each column value R(j) as the correction coefficient S(j) between samples, S(j) = Median[R(j)] Dividing the original data K(ij) by the coefficient column by column to obtain the corrected data K'(ij), K'(ij) = K(ij) / S(j) Wherein, the original data is a measurement value matrix K(ij) with N rows and M columns, where i = 1,...N represents proteins and j = 1,...M represents samples, and the samples include the data of two groups before and after treatment of each patient; K(i) is row data and K(j) is column data.

3. The treatment response data processing system based on the multiplex protein detection technology according to claim 1, wherein: In the process of gradually retaining the p-value data in the data by the feature variable selection module, the data of the first 100%, 50%, 25%, 10% and 5% of the data are sequentially selected for retention.

4. The treatment response data processing system based on the multiplex protein detection technology according to claim 3, wherein: The retention percentage value of the number of variable data in the feature variable selection module reduced to the original data quantity is 5%.

5. The treatment response data processing system based on the multiplex protein detection technology according to claim 2, wherein The method steps of principal component analysis in the feature variable selection module are as follows: Decomposing the processed data X into three matrices through singular value decomposition, X = USV T where U is the left singular vector matrix; S is the diagonal matrix of singular values; V T is the transpose of the right singular vector matrix; the data X is the corrected data K'(ij); Calculating the principal component scores, and the formula for calculating the principal component scores is, Score=U(n-1) 1 / 2 where n is the number of columns of U, and the first column Score i is the score of the i-th principal component.

6. The treatment response data processing system based on the multiplex protein detection technology according to claim 5, wherein The formula for calculating the change value of the first principal component is: Ds = |Score 1 [B]–Score 1 [A]| Wherein, Ds is the change value of the first principal component (Delta score, Ds); B is before treatment and A is after treatment.

7. The treatment response data processing system based on the multiplex protein detection technology according to claim 1, wherein The working steps of the subgroup evaluation module are as follows: Set the total number of patients as N and the number of subgroup patients as M; Sample the preprocessed data multiple times, with the number of patients drawn each time being M, to form random subgroups; Calculate the number of differentially expressed proteins before and after treatment within each subgroup drawn each time; Judge whether the number of differentially expressed proteins in the subgroup obtained through discrimination is greater than the number of differentially expressed proteins in 95% of the random subgroups. When the condition is met, it is judged that the subgroup obtained is reasonable.

8. The treatment response data processing system based on the multiplex protein detection technology according to claim 7, characterized in that: Calculate the number of differentially expressed proteins before and after treatment within each subgroup drawn each time by means of paired t-test.