Detection method combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay
By combining anti-NF155 antibody detection and enzyme-linked immunosorbent assays, combined with XGBoost model algorithm and autologous clinical data, early accurate diagnosis and typing of CIDP is achieved, and personalized treatment plans are provided, which solves the problems of diagnosis and treatment in the existing technology and improves the therapeutic effect and accuracy.
Patent Information
- Application Number
- CN202510274324.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-03
AI Technical Summary
It is difficult for the prior art to achieve accurate diagnosis and typing of chronic inflammatory demyelinated polyneuropathy (CIDP). Especially when dealing with complex disease manifestations, traditional methods cannot achieve comprehensive and accurate personalized diagnosis and treatment.
The detection method combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay was adopted, and an analytical model based on the XGBoost model algorithm was established, combined with the patient's autologous clinical data, CIDP subtype classification was performed, and personalized treatment plans were provided.
Accurate diagnosis and classification of CIDP is achieved, personalized treatment plans are provided, the accuracy and effectiveness of treatment are improved, and the treatment plans are timely adjusted through dynamic data analysis, improving the overall health management level of patients.
Smart Images

Figure CN120089402A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical detection, and particularly relates to a detection method combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay. Background Art
[0002] Chronic inflammatory demyelinating polyneuropathy (CIDP) is an immune-mediated neurological disease, mainly manifested as demyelination of peripheral nerves, resulting in gradually increasing muscle weakness, sensory loss, and motor dysfunction. The diagnosis of CIDP usually relies on clinical symptoms, detection of nerve conduction velocity, and other auxiliary laboratory tests. However, due to the multiple subtypes of CIDP and the similarity of its symptoms to other neuropathies, traditional diagnostic methods pose great challenges, especially in the early stage, where it is difficult to accurately diagnose and classify. Therefore, the accurate diagnosis and personalized treatment of CIDP have always been important research directions in the field of neurology.
[0003] In recent years, the discovery of anti-NF155 antibody has provided a new biomarker for the diagnosis of CIDP. Especially in certain subtypes, the increase in anti-NF155 antibody level is closely related to the pathogenesis of the disease. However, simple antibody detection and traditional clinical judgment methods are difficult to comprehensively analyze the disease characteristics of patients. Especially when dealing with complex disease manifestations, it is often impossible to achieve comprehensive and accurate personalized diagnosis and treatment.
[0004] According to the disclosed technical solutions, the technical solution with the publication number CN115094132A proposes an intestinal biomarker for evaluating chronic inflammatory demyelinating polyneuropathy and its application, and this biomarker can be obtained from the fecal samples of patients. The technical solution with the publication number US10509033B2 proposes a kit and biomarker for CIDP, which can perform a preliminary diagnosis related to CIDP on patients by detecting NF155 antibody. The technical solution with the publication number WO2016117618A1 proposes a diagnostic method for CIDP, which also determines the CIDP condition of patients by fluorescence labeling and measuring anti-NF155 antibody.
[0005] The above technical solutions all propose various diagnostic technical solutions for CIDP. With the in-depth research in related fields, more efficient detection methods and supporting detection systems can be further proposed.
[0006] The foregoing discussion of the background art is only intended to facilitate the understanding of the present invention. This discussion does not recognize or admit any part of the materials mentioned as common general knowledge. Summary of the Invention
[0007] The object of the present invention is to disclose a detection method combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay. The detection method uses the detection data of anti-NF155 antibody detection and enzyme-linked immunosorbent assay of patients, and combines with the autologous clinical data of patients to establish an analysis model based on the XGBoost model algorithm to classify the CIDP subtypes of patients, and provide personalized treatment plan suggestions in combination with the dynamic data of the patients' conditions. At the same time, a corresponding detection system is proposed, including a laboratory, a data repository and an analysis unit. The anti-NF155 antibody detection is performed by laboratory equipment, and the detection data and the autologous clinical data of the patients are stored in the data repository. The analysis unit establishes an analysis model based on the machine learning algorithm, and uses the detection data and the autologous clinical data of the patients as the input of the analysis model to classify the subtypes of the chronic inflammatory demyelinating polyneuropathy suffered by the patients, analyze the current dynamic data of the patients' conditions, and provide a reference treatment plan.
[0008] The present invention adopts the following technical solutions: A detection system combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay, the detection system includes: a laboratory, a data repository and an analysis unit; the data repository is communicatively connected with the laboratory and the analysis unit, and the detection data of the laboratory is stored in the data repository; and, based on the analysis requirements of the analysis unit, the data repository provides data to the analysis unit; The laboratory is configured to perform a plurality of detections and provide the obtained detection data; The data repository is configured to store detection data and also includes storing the autologous clinical data of multiple patients; The analysis unit is configured to establish an analysis model based on the machine learning algorithm, and use the detection data and the autologous clinical data of the patients as the input of the analysis model to classify the subtypes of the chronic inflammatory demyelinating polyneuropathy suffered by the patients, analyze the current dynamic data of the patients' conditions, and provide a reference treatment plan.
[0009] Preferably, the machine learning algorithm adopted by the analysis model is the XGBoost algorithm.
[0010] Preferably, the laboratory at least includes the equipment required for performing the enzyme-linked immunosorbent assay for anti-NF155 antibody detection.
[0011] At the same time, a detection method combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay is proposed, and the detection method is applied to the detection system; the detection method includes the following steps: S100: Generate a patient data set of the patient; the patient data set includes: the detection data of the patient's anti-NF155 antibody in the enzyme-linked immunosorbent assay and the autologous clinical data; S200: Perform data structuring processing on the patient data set and store it in the data repository; S300: Take the patient dataset of a patient as input, apply an analysis model based on machine learning algorithms to analyze the patient, output the subtype of the chronic inflammatory demyelinating polyneuropathy type suffered by the patient, analyze the current condition dynamic data of the patient, and provide a reference treatment plan.
[0012] Preferably, in step S100, the detection data includes one or more of the following data: antibody concentration, reaction kinetic curve, antibody subtype distribution, and experimental conditions; The autologous clinical data includes one or more of the following data: patient demographic information, past medical history and related diagnostic information, laboratory and molecular pathology test information, treatment data information, treatment outcome data information.
[0013] Preferably, the establishment of the analysis model based on machine learning algorithms includes the following steps: E100: Extraction and preprocessing of the original dataset to obtain a training dataset; E200: Label the training dataset, group it into a training set and a validation set, and perform feature selection and dimensionality reduction processing on the data; E300: Perform model training using the training set based on the XGBoost model algorithm; E400: Use the validation set to verify the performance and accuracy of the established model, and finally establish the analysis model.
[0014] Preferably, in step E200, at least one of the principal component analysis method or the LASSO regression algorithm is used to perform feature selection and dimensionality reduction processing on the data.
[0015] The beneficial effects achieved by the present invention are: Through the combination of anti-NF155 antibody detection and enzyme-linked immunosorbent assay technology in this technical solution, this solution can provide accurate diagnosis and typing of chronic inflammatory demyelinating polyneuropathy for patients. With the help of a machine learning model, different subtypes of CIDP can be classified based on antibody concentration, clinical information, and other experimental data, thus helping doctors accurately judge the disease type of patients and providing data support for subsequent treatment; This technical solution not only classifies CIDP, but also generates personalized treatment recommendations based on the detection results and autologous clinical data of patients, combined with the existing treatment database. The model can dynamically recommend suitable treatment plans for patients, including treatment courses, medication dosages, treatment methods, etc., based on the reaction of anti-NF155 antibodies, treatment response history, and genomic characteristics of patients, thereby improving the accuracy and effectiveness of treatment; This technical solution can integrate the patient's autologous clinical data, anti-NF155 antibody test data, and other laboratory test results through a data repository, and use machine learning algorithms to perform dynamic data analysis on the patient. As the treatment progresses, the patient's disease changes and treatment responses can be fed back to the analysis unit in real time, providing accurate dynamic trend analysis for doctors. This data integration and dynamic monitoring helps to adjust the treatment plan in a timely manner, ensure that the patient receives the most appropriate treatment, and thus improve the overall health management level of the patient. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention can be further understood from the following description in conjunction with the accompanying drawings. The components in the drawings are not necessarily drawn to scale, but the emphasis is placed on showing the principles of the embodiments. In different views, the same reference numerals designate corresponding parts.
[0017] Serial number description: 101 - detection system; 102 - patient; 103 - test result data; 104 - autologous clinical data; 115 - sample; 140 - laboratory; 160 - data repository; 180 - analysis unit; 182 - analysis model; 190 - output; 500 - computer system; 502 - bus; 504 - processor; 506 - main memory; 508 - read-only memory; 510 - storage device; 512 - display; 514 - input device; 516 - cursor control device; 518 - network device; Figure 1 Schematic diagram of the detection system described in the embodiment of the present invention Figure 2 Schematic diagram of the steps of the detection method described in the embodiment of the present invention; Figure 3 Schematic diagram of the training steps of the analysis model described in the embodiment of the present invention; Figure 4 Schematic diagram of the computer system adopted by the system in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with its embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. For those skilled in the art, after referring to the following detailed description, other systems, methods, and / or features of this embodiment will become obvious. It is intended that all such additional systems, methods, features, and advantages are included in this specification. Included within the scope of the present invention and protected by the appended claims. Additional features of the disclosed embodiments are described in the following detailed description, and these features will be obvious based on the following detailed description.
[0019] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific orientation and be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and cannot be construed as a limitation of this patent. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0020] Example 1: Exemplarily, as shown in the attached Figure 1 figure, a detection system combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay is proposed. The detection system includes: a laboratory, a data repository, and an analysis unit; the data repository is communicatively connected to the laboratory and the analysis unit, and the detection data of the laboratory is stored in the data repository; and, based on the analysis requirements of the analysis unit, the data repository provides data to the analysis unit. The laboratory is configured to perform multiple detections and provide the obtained detection data. The data repository is configured to store detection data and also includes storing the autologous clinical data of multiple patients. The analysis unit is configured to establish an analysis model based on a machine learning algorithm, use the detection data and autologous clinical data of the patient as the input of the analysis model, classify each subtype of the chronic inflammatory demyelinating polyneuropathy type suffered by the patient, analyze the current condition dynamic data of the patient, and provide a reference treatment plan.
[0021] Preferably, the machine learning algorithm adopted by the analysis model is the XGBoost algorithm.
[0022] Preferably, the laboratory at least includes the equipment required for performing the enzyme-linked immunosorbent assay for anti-NF155 antibody detection.
[0023] Meanwhile, a detection method combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay is proposed. The detection method is applied to the detection system; as shown in the attached Figure 2 figure, the detection method includes the following steps: S100: Generate a patient dataset for the patient; the patient dataset includes: the detection data of the patient's anti-NF155 antibody in the enzyme-linked immunosorbent assay and the autologous clinical data. S200: Perform data structuring processing on the patient dataset and store it in the data repository. S300: Take the patient data set of a patient as input, apply an analysis model based on a machine learning algorithm to analyze the patient, output the subtype of the chronic inflammatory demyelinating polyneuropathy type suffered by the patient, analyze the current condition dynamic data of the patient, and provide a reference treatment plan.
[0024] Preferably, in step S100, the detection data includes one or more of the following data: antibody concentration, reaction kinetics curve, antibody subtype distribution, and experimental conditions; The autologous clinical data includes one or more of the following data: the patient's demographic information, past medical history, and relevant diagnostic information, laboratory and molecular pathology test information, treatment data information, treatment outcome data information.
[0025] Specifically, referring back to the appendix Figure 1 , shows an implementation architecture of the detection system in an exemplary embodiment. Among them, the detection system 101 may include a laboratory 140, a data repository 160, and an analysis unit 180. The data repository 160 is communicatively connected to the laboratory 140 and the analysis unit 180 so that the results of the laboratory 140 can be stored in the data repository 160. And, based on the analysis requirements of the analysis unit 180, the data repository 160 can provide various information sets to the analysis unit 180. The analysis unit 180 has at least one input for receiving the results or other information from the laboratory 140 and generates at least one output 190.
[0026] In a preferred embodiment, the autologous clinical data 104 of each patient 102 of the detection system 101 is established as a file and stored in the data repository 160. The autologous clinical data 104 can be generated according to various clinical information in the patient's individual health record. The clinical information can be summarized and recorded based on the information entered into the electronic medical record or electronic health record by doctors, nurses, or other medical professionals.
[0027] The autologous clinical data 104 also includes information of the patient 102 at various time nodes in various states. Preferably, the autologous clinical data 104 can also select appropriate information according to the type of disease suffered by the patient to streamline the data volume of the autologous clinical data.
[0028] For example, for cancer patients, collecting clinical information in the field of cancer can include: demographic information such as year of birth, gender, race / ethnicity, related comorbidities, smoking history; diagnostic information such as sample tissue source, date of initial diagnosis, histology, grade, metastatic diagnosis, metastatic site, TNM stage, etc.; laboratory and molecular pathology test information such as laboratory type, test results and units, event date, etc.; treatment data information such as drug name, start and end dates, dose, number of cycles, type and date of surgery, radiotherapy location and mode, etc.; treatment outcome data information such as efficacy response, date of disease progression or recurrence, adverse reactions, etc.
[0029] And preferably, the autologous clinical data 104 can further include genomic feature data of the patient 102; the genomic feature data can be obtained from RNA or DNA sequencing; where gene variations can be analyzed, for example, including performing single nucleotide polymorphism (SNP) detection, insertion / deletion events, loss-of-function / gain-of-function detection, fusion events, copy number variation calculation, microsatellite instability assessment, and other structural variations in DNA and RNA.
[0030] In a preferred embodiment, the laboratory 140 is capable of fully performing an enzyme-linked immunosorbent assay for detecting anti-NF155 antibodies. The laboratory 140 is preferably configured with corresponding dedicated and general equipment to meet the needs of sample processing, analysis, and result output.
[0031] Preferably, the laboratory 140 is equipped with basic equipment for sample preparation, processing, storage, and transportation of the sample 115 of the patient 102, including glass slides, microscopes, graduated instruments, glassware, plastic products, tweezers, scalpels, blood coagulation tubes, vacuum tubes, freezer bags, temperature-controlled storage equipment, and similar devices. In addition, it also includes professional equipment for analysis, such as mass spectrometers, chromatographs, titrators, spectrometers, particle analyzers, and rheometers, etc. These devices can meet various basic needs during the experimental process.
[0032] Preferably, the laboratory 140 is configured with equipment related to enzyme-linked immunosorbent assay, including but not limited to microplate washers, centrifuges, microplate readers, refrigeration equipment, pipettes and tips, dispensers, thermostatic incubation shakers, and microplates, etc. In addition, it also includes photometers, spectrophotometers, reagent dispensers, and related data analysis software for achieving high-precision detection. The laboratory provides standardized ELISA kits and can also develop customized experiments according to requirements.
[0033] Preferably, the laboratory 140 can also be configured with other detection equipment to assist in the supporting detection after the enzyme-linked immunosorbent assay, such as next-generation sequencing (NGS) equipment, imaging and pathology equipment, or organ-class laboratory support equipment, etc.
[0034] Preferably, through the laboratory 140, the test result data 103 of multiple test indicators of the patient 102 can be obtained. Preferably, the test result data 103 includes one or more of the following test values: (1) Anti-NF155 antibody concentration (OD value); the OD value can reflect the level of antibody concentration in the sample, and different CIDP subtypes may show significantly different antibody levels; a high concentration of anti-NF155 antibody is usually associated with the NF155-positive CIDP subtype.
[0035] (2) Detection of IgG subtype ratio, where whether the ratio of IgG4 is significantly increased is an important feature of NF155-positive CIDP. It is calculated by the following method: IgG4 ratio = OD value of IgG4 / OD value of total IgG.
[0036] For NF155-positive CIDP patients, the ratio of IgG4 is usually significantly increased.
[0037] Dilution factor curve; the change in antibody concentration at different dilution factors can reflect the affinity of antigen-antibody binding. By plotting the curve of OD value versus dilution factor, the half-maximal inhibitory concentration (IC50), that is, the dilution factor when the antibody concentration is reduced by half, can be extracted to reflect the antibody affinity. Different NF155-positive patients may have different antibody affinity characteristics for different subtypes.
[0038] Reaction kinetic data, which characterizes the change of OD value with reaction time and can reveal the antigen-antibody binding rate. This data includes the initial rate and the plateau time; through this data, it can be identified that NF155-positive CIDP may have specific binding kinetic characteristics.
[0039] Further, it may also include non-specific binding signals, which are used to characterize the OD values of blank wells and non-specific antibody wells and reflect the degree of interference of the detection background; detection conditions, such as reaction temperature and incubation time and other values, to assist in the establishment of the analysis model.
[0040] In a preferred embodiment, the data repository 160 includes data structure transformation and processing of the input data, so that the processed data is conducive to further storage and use. Preferably, the input data undergoes the following three processing stages: (1) Original dataset; the original dataset stores the original forms of the detection result data 103 and the autologous clinical data 104. The data in the original form may include various forms such as scanned copies of handwritten health records, genomic test reports, other laboratory reports, drug lists, image files, etc. For example, for the detection data of anti-NF155 antibodies, these data may include the original records of ELISA tests and relevant information in the clinical history (such as nerve conduction velocity, cerebrospinal fluid protein concentration, etc.). These records may come from multiple medical institutions or doctors where the patients seek medical treatment, and the written records are digitized through optical character recognition (OCR) technology and stored in a structured file format.
[0041] (2) Structured dataset; the structured dataset is obtained by extracting information from the original dataset and standardizing it, converting the patient's health information into a structured form with tags or descriptors. Preferably, the data storage format of the structured dataset includes tabular form, relational database, and other formats suitable for medical records, such as image files, binary files, etc.
[0042] Preferably, the data content of the structured dataset: Treatment information: such as the immunotherapy regimens for CIDP patients (such as glucocorticoids, intravenous immunoglobulin (IVIG), or rituximab).
[0043] Laboratory data: including ELISA test data (OD values, IgG subtype distribution), dilution curve characteristics, kinetic parameters, etc.
[0044] Patient characteristics: such as age, gender, occupation, blood type, BMI, etc.
[0045] Monitoring data: such as blood glucose, blood pressure, weight, and other long-term health monitoring parameters.
[0046] Medication data: including drug name, dosage, medication time, prescription date, etc.
[0047] And preferably, the data repository 160 also includes a structured clinical big data repository. The structured datasets of multiple patients are stored in the clinical big data repository, providing support for large-scale analysis and the training of machine learning models.
[0048] Preferably, the data scale of the data repository 160 can cover the data of thousands to millions of patients.
[0049] Preferably, the data in the data repository 160 can be de-identified to ensure patient privacy while retaining the effectiveness of data analysis.
[0050] Preferably, the data in the data repository 160 is normalized, and specific health information is mapped by numerical identifiers by using SNOMED equivalent sets.
[0051] In a preferred embodiment, the analysis unit 180 extracts the required data from the data repository 160 for training based on a machine learning algorithm, and establishes an analysis model 182. As the core part of the detection system 101, it will be described later.
[0052] In a preferred embodiment, the output 190 may include the following: Diagnosis and classification results of CIDP: Based on anti-NF155 antibody detection and patient clinical information, provide CIDP diagnosis suggestions and disease classification; Dynamic disease data analysis: Integrate laboratory results and clinical data to identify potential disease patterns or trends; Treatment plan recommendation: Combine known treatment databases or antibody response data to generate individualized treatment suggestions, such as the length of treatment cycle, dosage, treatment methods, etc.
[0053] Embodiment 2: This embodiment should be understood as including at least all the features of any of the foregoing embodiments, and further improved on this basis: In a preferred embodiment, as shown in the appendix Figure 3 The establishment and training of the analysis model 182 includes the following steps: E100: Data extraction and preprocessing; The analysis unit 180 extracts the data required for model training from the data repository 160 as the original data set. Preferably, in the preprocessing step, the following processes are included: (1) Remove or fill in missing values, inconsistent values, and outliers. (2) Normalize numerical data, such as normalization of values. (3) Encode categorical variables, such as gender (0: male, 1: female).
[0054] Furthermore, feature engineering is performed on non-direct data to calculate key indicators. Such indicators include the IgG4 ratio of anti-NF155 antibody, half-maximal inhibitory concentration, kinetic parameters, etc.
[0055] E200: Data refinement and dimensionality reduction. Among them, it further includes the following sub-steps: E210: Label the data, including: Classify the IDP subtype labels: Based on anti-NF155 antibody detection results, clinical diagnosis, etc., CIDP can be divided into different subtypes, such as typical CIDP, CIDP subtypes related to positive anti-NF155 antibody, etc.
[0056] Treatment response label: Samples are labeled according to the patient's treatment response, such as efficacy score, dosage, treatment cycle, cycle recovery, etc., to facilitate subsequent treatment plan recommendations.
[0057] E220: Data splitting. The obtained overall dataset is divided into a training set and a validation set according to the allocation ratio of 80% and 20%.
[0058] E230: Feature selection and dimensionality reduction processing are performed on the data.
[0059] Preferably, the principal component analysis method can be used to process the data to reduce data complexity, remove redundant information, and improve the performance of the machine learning model.
[0060] Specifically, the principal component analysis method projects the original high-dimensional data into a low-dimensional space through a linear transformation while retaining the main variation information of the data. Assume the original data matrix is X Î R n*p , where R is the set of real numbers, n is the number of samples, and p represents the number of features. The goal of PCA is to project X into a low-dimensional space through the feature covariance matrix to obtain the dimensionality-reduced dataset Z, Z Î R n*k , and k ≪ p.
[0061] Furthermore, in the detection data, multiple features may have significant correlations. For example, antibody concentration and IgG subtype ratio, characteristic points of the dilution multiple curve and kinetic parameters, etc. The redundancy between these features will increase the complexity of the model. PCA removes redundant information by selecting the main feature vectors to generate uncorrelated new features. Preferably, the main components with a cumulative explained variance ratio reaching 95% can be selected to retain the main information in the data.
[0062] Furthermore, the LASSO (Least Absolute Shrinkage and Selection Operator) regression can also be used to evaluate the importance of features and retain the most diagnostically valuable features. By applying the LASSO regression method to evaluate the impact of each feature on the diagnostic result and effectively retain the most diagnostically valuable features. The core purpose of this process is to screen out the features most relevant to disease diagnosis from the original data, thereby improving the accuracy and generalization ability of the prediction model.
[0063] Among them, the objective function of LASSO regression can be expressed as: ; In the above formula, y i is the observed value of the i-th sample, x ij is the j-th feature of the i-th sample, β j$\beta_j$ is the regression coefficient of the $j$-th feature in the regression model, and $\lambda$ is the regularization parameter that controls the shrinkage intensity of the feature coefficients. The key of LASSO regression lies in controlling the strictness of feature selection by adjusting the value of $\lambda$. When $\lambda$ is large, the model tends to compress the coefficients of more features to zero, thus selecting fewer most diagnostically valuable features; when $\lambda$ is small, the model will retain more features. Preferably, by performing LASSO regression on each feature in the dataset, the regression coefficients $\beta_j$ of each feature can be obtained. j The numerical value. By the magnitudes of these coefficients, combined with the selection of the regularization parameter $\lambda$, the contribution degree of each feature to the target variable can be quantified. In actual operation, a larger regression coefficient $\beta_j$ j indicates that this feature has a greater impact on the prediction result of the model; while a feature with a coefficient of zero indicates that this feature has a smaller predictive effect on the target variable under the current dataset and can be excluded.
[0064] One or both of the above principal component analysis method or LASSO regression algorithm can be selected. This depends on the number of features and the amount of data in the original data. When the number of features in the original data is too large and there is a high correlation between features (a lot of redundant information), the principal component analysis method can map the features onto a small number of principal components while retaining the main variation information of the data. When it is necessary to clearly evaluate the importance of each feature and the number of features is moderate, LASSO regression can directly perform screening.
[0065] E300: Model training. In a preferred embodiment, the XGBoost model is selected as the machine learning algorithm model of the analysis model.
[0066] Among them, the XGBoost model is an algorithm model based on the gradient boosting tree algorithm. XGBoost performs classification or regression tasks by constructing multiple decision trees and combining them. The construction of the decision tree includes: Split selection, that is, dividing the dataset into two subsets by different values of the feature to find the best split point (such as the Gini index or information gain) at each node; Leaf node prediction, each leaf node of the decision tree will contain a predicted value.
[0067] Preferably, the XGBoost model does not generate a tree all at once, but constructs trees through multiple iterations. Each new tree attempts to correct the previous model by reducing the error of the previous tree to achieve the purpose of training. In each round of iteration, the newly generated tree is determined by calculating the gradient. The steps of gradient boosting can be summarized as follows: in each iteration, calculate the gradient of the loss function with respect to the output of the existing model; generate a tree that minimizes this gradient as much as possible; add the newly generated tree to the existing model to generate a new combined prediction model. Repeat the above process until a predetermined number of trees or other stopping criteria are reached. Therefore, in XGBoost, the construction of each tree is actually to correct the prediction error of the current model and gradually improve the accuracy of the model. Further, the model training includes the following sub-steps: E310: Define the objective function Target(θ): ; In the above formula, is the loss function, representing the error between the predicted value and the actual value of each sample, and y i is the true value of the i-th sample; is the predicted value of the model.
[0068] is the regularization term, which is used to prevent overfitting during training. The form of the regularization term is the complexity of the policy tree, usually using L1 or L2 regularization. The regularization term is used for the split selection of each tree to control the depth of the tree and the number of splits. can be expressed as: ; In the above formula, T is the number of leaf nodes of the current decision tree; γ is the penalty parameter for tree splitting, used to control the growth of the number of leaf nodes; λ is the regularization parameter, which controls the complexity of each node and is used to control the sum of the squares of the weights of each leaf node to prevent overfitting.
[0069] Using the gradient boosting algorithm, it depends on training a new tree based on the residual information of the previous step. The goal of each new tree is to correct the error of the previous tree. Assume that in the t-th step of training, the model prediction is , and the goal is to update the model by optimizing the residual: ; In the above formula, is the current predicted value; η is the learning rate, preferably η is less than 1; p t+1 (x) is the newly trained decision tree.
[0070] Minimize the objective function by continuously splitting nodes. The splitting criterion can be based on information gain or the Gini index. Preferably, a splitting criterion based on quadratic approximation can be adopted. Specifically, the gain calculation formula during splitting is: ; In the above formula, g i and h i respectively represent the elements of the gradient Hessian matrix, reflecting the error and its rate of change of each data point; L and R respectively represent the left and right child nodes.
[0071] Finally, the final predicted value is obtained by weighting the outputs of all trees: ; In the above formula, p k (x) is the output of the k-th decision tree.
[0072] E320: Set the following hyperparameters and tune each hyperparameter. The hyperparameters include: The depth of the tree, which controls the complexity of the tree and avoids overfitting; The learning rate, which controls the update amplitude of each step; The subsample ratio, randomly select part of the samples for each round of training to improve the generalization ability of the model; Column sampling, the number of features randomly selected during the training of each decision tree, reducing overfitting.
[0073] E330: Perform cross-validation on the established model. Preferably, K-fold cross-validation is adopted to evaluate the performance of the model and ensure the stability and generalization ability of the model.
[0074] E400: After training is completed, evaluate the performance of the XGBoost model through the validation set. Preferably, the evaluation metrics include the accuracy of model classification, the AUC, i.e., the area under the ROC curve, to evaluate the discrimination ability of the classifier; optionally, it also includes analyzing the precision, recall rate, and F1 score of the model.
[0075] Furthermore, apply the established analysis model as the main model and further train multiple submodels to improve the treatment plan for patients. Among them, the submodels can include: Drug response model: Use the drug response simulation model to analyze the efficacy of different drugs in different CIDP subtypes. This model can combine pharmacological data, such as the pharmacokinetics / pharmacodynamics data of drugs and the individual differences of patients, to simulate and predict the possible drug responses of patients.
[0076] Drug - Gene Interaction Model: Based on the patient's genomic data, by simulating the interaction between drugs and genes, it predicts the efficacy of drug treatment and possible side effects. Combining gene variations (such as SNP analysis) with drug efficacy data helps to determine individualized treatment plans.
[0077] Clinical Treatment Pathway Model: A pathway model can be established based on historical patient treatment data to simulate the disease progression trends under different treatment paths. By evaluating the efficacy at different time points during the treatment process, it recommends the most suitable treatment cycle and drug combination for the current patient.
[0078] Preferably, the classification results of the main model of the analysis model are combined with the drug response model, the drug - gene interaction model, and the clinical pathway simulation to form a comprehensive treatment recommendation framework. Based on different variables output by the model (such as CIDP subtype, drug response, gene variation), personalized treatment plans are generated.
[0079] Specifically, the treatment plan may include: Length of Treatment Course: The model analyzes the patient's CIDP subtype and historical treatment data, combined with drug response prediction, to recommend the most suitable treatment course for the patient, such as whether long - term immunotherapy is required or a shorter - cycle treatment plan.
[0080] Dosage and Drug Selection: Based on the analysis model to judge the patient subtype and drug response simulation, individualized drug dosages and types are recommended. For example, for anti - NF155 - positive CIDP patients, immunoglobulin (IVIG) and glucocorticoids may be more preferred, but the treatment dosages and durations may vary among different patients.
[0081] Adjuvant Treatment Measures: Such as physical therapy, rehabilitation training, nutritional support, etc., can also provide supplementary treatment measures according to the recommended path of the model.
[0082] Example 3: This example should be understood as including at least all the features of any of the foregoing examples and being further improved on that basis: Exemplarily, attached Figure 4 depicts a schematic diagram of a computer system 500 in which the analysis unit 180 described herein can be implemented; the computer system 500 can, according to the current control program, achieve the work control of each working unit, module, and component in the system; and, it also includes collecting, storing, and processing the work data and detection data generated during the system operation to ultimately achieve the expected effect of the fluorescence labeling system.
[0083] Among them, computer system 500 includes bus 502 or other communication mechanisms for transmitting information, and one or more processors 504 coupled to bus 502 for processing information; the processor 504 can be, for example, one or more general-purpose microprocessors; Computer system 500 also includes main memory 506, such as random access memory (RAM), cache, and / or other dynamic storage devices, which are coupled to bus 502 for storing information and instructions to be executed by processor 504; main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions executed by processor 504; when these instructions are stored in a storage medium accessible by processor 504, computer system 500 is presented as a special-purpose machine customized to execute the operations specified in the instructions; Computer system 500 may also include read-only memory (ROM) 508 or other static storage devices coupled to bus 502 for storing static information and instructions of processor 504; a storage device 510 such as a magnetic disk, optical disk, or USB drive (flash drive) will be coupled to bus 502 for storing information and instructions; And further, coupled to bus 502 may also include a display 512 for displaying various information, data, media, etc., and an input device 514 for allowing a user of computer system 500 to control, manipulate, and / or interact with computer system 500; A preferred way to interact with the management system can be through a cursor control device 516, such as a computer mouse or similar control / navigation mechanism; Furthermore, computer system 500 may also include a network device 518 coupled to bus 502; where the network device 518 may include components such as a wired network card, wireless network card, switching chip, router, switch, etc.; Generally, the terms "engine", "component", "system", "database", etc. used herein may refer to logic embodied in hardware or firmware, or to a collection of software instructions, which may have entries and exit points and are written in a programming language such as Java, C, or C++; software components can be compiled and linked into an executable program, installed in a dynamic link library, or can be written in an interpreted programming language (such as BASIC, Perl, or Python); it should be understood that software components can call other components or themselves, and / or can be called in response to detected events or interrupts; A software component configured to execute on a computing device can be provided on a computer-readable medium, such as a compact disc, digital video disc, flash drive, magnetic disk, or any other tangible medium, or as a digital download (and can initially be stored) in a compressed or installable format that requires installation, decompression, or decryption before execution); such software code can be stored, in whole or in part, on the memory device of the executing computing device for execution by the computing device; the software instructions can be embedded in firmware, such as an EPROM. It should also be understood that hardware components can consist of connected logic units (such as gates and flip-flops), and / or can consist of programmable units (such as programmable gate arrays or processors); The computer system 500 can include custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic to implement the techniques described herein, and the program logic, in combination with the computer system, causes the computer system 500 to be a special-purpose computing device; In accordance with one or more embodiments, the techniques herein are performed by the computer system 500 in response to one or more sequences of one or more instructions contained in the main memory 506 being executed by the processor 504; such instructions can be read from another storage medium, such as the storage device 510, into the main memory 506; the execution of the sequence of instructions contained in the main memory 506 causes the processor 504 to perform the processing steps described herein; in an alternative embodiment, hardwired circuitry can be used in place of or in combination with software instructions; As used herein, the term "non-transitory medium" and like terms refer to any medium that stores data and / or instructions that cause a machine to operate in a particular manner; such non-transitory media can include non-volatile media and / or volatile media; non-volatile media includes, for example, optical discs or magnetic disks, such as the storage device 510; volatile media includes dynamic memory, such as the main memory 506; Common forms of non-transitory media include, for example, floppy disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with a hole pattern, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, and network versions thereof; Non-transient media is different from transmission media but can be used in combination with transmission media; transmission media participates in the transfer of information between non-transient media; for example, transmission media includes coaxial cables, copper wire, and optical fiber, including the wires that make up the bus 502; transmission media can also take the form of acoustic or light waves, such as radio waves and infrared data communication.
[0084] Although the present application has been described above with reference to various embodiments, it should be understood that many changes and modifications can be made without departing from the scope of the present application. That is to say, the methods, systems and devices discussed above are examples. Various configurations can be appropriately omitted, replaced or added with various processes or components. For example, in an alternative configuration, the methods can be performed in an order different from the described order, and / or various components can be added, omitted and / or combined. Moreover, the features described with respect to certain configurations can be combined in various other configurations, such as different aspects and elements of the configurations can be combined in a similar manner. In addition, as technology develops, the elements therein can be updated, that is, many elements are examples and do not limit the scope of the present disclosure or the claims.
[0085] Specific details are given in the description to provide a thorough understanding of the exemplary configurations including the implementation. However, the configurations can be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures and technologies have been shown without unnecessary details to avoid obscuring the configurations. The description only provides example configurations and does not limit the scope, applicability or configuration of the claims. On the contrary, the foregoing description of the configurations will provide those skilled in the art with an enabling description for implementing the described technology. Various changes can be made to the functions and arrangements of the elements without departing from the spirit or scope of the present disclosure.
[0086] In summary, it is intended that the above detailed description be considered illustrative rather than restrictive, and it should be understood that the above embodiments should be construed as only for illustrating the present invention and not for limiting the protection scope of the present invention. After reading the content recorded in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A detection system combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay, characterized in that: The detection system comprises: a laboratory, a data storage repository and an analysis unit; the data storage repository is in communication with the laboratory and the analysis unit, and the detection data of the laboratory is stored in the data storage repository; and based on the analysis requirements of the analysis unit, the data storage repository provides data to the analysis unit; The laboratory is configured to perform a plurality of tests and provide the obtained test data; The data repository is configured to store test data and also includes storing autologous clinical data of a plurality of patients; The analysis unit is configured to establish an analysis model based on a machine learning algorithm, and use the patient's test data and autologous clinical data as input to the analysis model, classify the patient's chronic inflammatory demyelinating polyneuropathy into subtypes, analyze the patient's current disease dynamics data, and provide a reference treatment plan.
2. The detection system according to claim 1, characterized in that: The machine learning algorithm used in the analysis model is the XGBoost model algorithm.
3. The detection system according to claim 2, characterized in that: The laboratory includes at least the equipment required for performing an ELISA for anti-NF155 antibody detection.
4. A detection method combining anti-NF155 antibody detection and enzyme-linked immunosorbent assay, characterized in that: The detection method is applied to the detection system as claimed in claim 3; the detection method comprises the following steps: S100: Generate a patient data set of the patient; the patient data set includes: detection data of the patient's anti-NF155 antibody in an enzyme-linked immunosorbent assay and autologous clinical data; S200: Perform data structuring on the patient data set and store it in a data repository; S300: Take a patient's patient data set as input, apply an analysis model based on a machine learning algorithm, analyze the patient, output the subtype of the chronic inflammatory demyelinating polyneuropathy type suffered by the patient, analyze the patient's current disease dynamics data, and provide a reference treatment plan.
5. The detection method according to claim 4, characterized in that: In step S100, the detection data includes one or more of the following data: antibody concentration, reaction kinetic curve, antibody subtype distribution and experimental conditions; The autologous clinical data includes one or more of the following data: patient's demographic information, past medical history and related diagnosis information, laboratory and molecular pathology test information, treatment data information, and treatment outcome data information.
6. The detection method according to claim 5, characterized in that: The establishment of the analysis model based on the machine learning algorithm includes the following steps: E100: Extraction and preprocessing of original data sets to obtain training data sets; E200: labeling the training data set, grouping it into a training set and a validation set, and performing feature selection and dimensionality reduction processing on the training data set; E300: Performing model training based on the XGBoost model algorithm using the training set; E400: Use the validation set to verify the performance and accuracy of the established model, and finally establish the analysis model.
7. The detection method according to claim 6, characterized in that: In step E200, at least one of a principal component analysis method and a LASSO regression algorithm is used to perform feature selection and dimensionality reduction processing on the data.
Citation Information
Patent Citations
Intestinal biomarker for evaluating chronic inflammatory demyelinating multiple nerve root neuropathy and application of intestinal biomarker
CN115094132A
Method, kit and biomarker for diagnosing chronic inflammatory demyelinating polyneuropathy
US10509033B2
Chronic inflammatory demyelinating polyneuropathy diagnostic method, kit, and biomarker
WO2016117618A1
TL1a patient selection methods, systems and devices
CN114375339A
Construction method and system of XGBoost machine learning model for judging autoimmune encephalitis prognosis
CN118039173A
Cited By
Anti-NF155 antibody detection material and preparation method thereof
CN122128366A