Method for predicting the effectiveness of chemotherapy treatment

The method uses serum N-glycome analysis and ML to predict chemotherapy efficacy in lung cancer by correlating structural glycan changes with treatment outcomes, addressing the lack of precision in existing methods and improving treatment effectiveness.

WO2025238388A1PCT designated stage Publication Date: 2025-11-20PANNON EGYETEM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/HU2025/050029
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-12
Filing Date
2025-05-12
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Current methods lack high-precision models to predict the effectiveness of chemotherapy in lung cancer patients, with conventional approaches relying on CT scans and MRIs, which are inadequate for determining chemotherapy efficacy, leading to ineffective treatments in approximately 70-80% of cases and significant side effects.

Method used

A method utilizing serum N-glycome analysis combined with machine learning (ML) to correlate structural changes in N-glycans with chemotherapy outcomes, employing capillary electrophoresis and laser-induced fluorescent detection to predict treatment efficacy through binary classification tasks.

Benefits of technology

Provides a precise and rapid tool for predicting chemotherapy effectiveness in lung cancer patients, achieving high discrimination power with ROC values greater than 0.9, enabling informed therapeutic decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000015_0001
    Figure IMGF000015_0001
  • Figure IMGF000016_0001
    Figure IMGF000016_0001
  • Figure IMGF000017_0001
    Figure IMGF000017_0001
Patent Text Reader

Abstract

The present invention relates to a method for predicting the effectiveness of chemotherapy treatment in lung cancer utilizing serum N-glycome analysis. The method enables an early detection whether the chemotherapy started in a patient diagnosed with lung cancer is really effective. Early detection of ineffective therapy is beneficial for the patient, as it is possible to change to another therapy in time.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR PREDICTING THE EFFECTIVENESS OF CHEMOTHERAPY TREATMENTFIELD OF THE INVENTION The present invention relates to a method for pcrtiendgi the effectiveness of chemotherapy treatment in lung cancer utilizing serum N-glycome analysis. BACKGROUND OF THE INVENTION Lung cancer represents a significant global he cahltahllenge, ranking as the second most prevalent malignant tumor worldwide [Bade, 202]0[Nooreldeen 202]1. The incidence of the disease is notably highin Hungary, where 10,600 cases were diagnosed2 in2 [ 2B0ogos in print, 202]4. On a European scale, theissue is staggering, with 484,000 cases reporte thde in same yea [rHarðardottir 202]2 and this number has been increasing since. Considering that lung ca insc tehre leading cause of malignancy-related de dauthes to its high incidence, late-stage discovery, comxp elteiology, heterogeneity and aggressive na [tSucrehabath 2019], more targeted and personalized approaches a hreigh in demand. Certain forms of lung cancer progress quickly, creating an urgent need to s etfafertctive treatment [sAlghamdi 2018]. Cancer treatment includes various types of interventions such agse sruyr, radiation therapy, immunotherapy, targetedra tphy, and several types of chemothera [Hpyirsch 2017]. A careful monitoring of the patients during can tcheerrapy is of high importance. Mészáros et al. developed an analytical methodh w chaicn be readily used to follow up or monitor the tumor surgery by investigating the glycan (sr)ug staructures of proteins. N-glycosylation pattern of serum samples of lunngc cear patients were analyzed before and after thesurgery using capillary electrophoresis separa wtioithn laser-induced fluorescent detection. The i rveelatpeak areas of N-glycans were evaluated from thueir aecdq electropherograms using machine learning-dbase data analysis. The study clearly aimed at the evaluation of thleati roenship between the clinical parameters(anamnesis), the result of operation and thev rela ptei ak area changes in the individual N-glycaunc stutr es,and correlations of the relative peak area altoenrasti of N-glycan groups with the clinical paramet weersre also found. Correlations of clinical parametersh w citertain glycome peaks varied. Several clinicalparameters were included in the study. Howevery, on el parameter (positive outcome of the surgedry ansmoker) from the 51 caused a change with satisf aycincguracy in the relative peak area of neutrala gnlysc. In general, this paper relates to monitoring thteie pnats’ status after tumor surgery and no idea of decision relating to therapy is or can be broug Whht.ile the authors planned to apply their workflo inw, which anamnesis of patients play an important ro nle m, onitoring the effects of chemotherapy on th-e N glycan profile, they are silent both on the expbelceta success of such a study or other result besides monitoring. However, an operation is a physicaelr ivnetntion while chemotherapy is chemical, and th isere no hint of any relationship between the two. Despite advances in oncology therapies, chemotyhe irsap still one of the most commonly used treatment modalities. Cytotoxic agents are used bo fotrh palliative and adjuvant purposes. Biomarkers predicting the effectiveness of chemotherapy wo bueld of great clinical importance, since chemotherapy proves to be ineffective in approximately 70-80% c oafses [Rosell 2006][Willers 2013]. Inadequately chosen chemotherapy drugs not only compromiseo thsseib pility of recovery but they can have serioudse si effects and also represent an economic burdene on a tthional health care syste[mYasng, K.A.-O 2023]. In fact, oncologists and healthcare professionhaolsul sd carefully plan and monitor the chemotherapytreatment outcome to minimize side effects whilfeici ef ntly targeting the tumo [rNagasaka 201]8. Thesetreatment plans are mainly dependent on canceryp seusbt and possibly existing genetic mutations. Therefore, it is particularly important to undernsdta the effectiveness of the chemotherapy treatm afetenrt the first session. Conventionally, clinicians usually test the effincciey of chemotherapy based on CT scans, and MRIsto make classifications as complete / partial rensp eo (regression), stable (stationary), or diseraosgere pssion(progression). While monitoring chemotherapy in one was or theer ot ish known in the field, there is an unmet need for high precision models to predict the effectievsesn of chemotherapy (EoC) of lung cancer patie Snutcsh. a predicting method could be a useful tool for c thlienician in her / his decision whether to continue chemotherapy or shift to an other treatment. Monitoring methods are conceptually different anod n dot help in themselves in prediction. Principally, there are three different ways to picrted EoC by utilizing 1) anamnesis containing alltraditional clinically relevant tests, 2) mono- a mndulti-omics data, and 3) based on imaging of , ce tislls uesor organs. In all these three cases, artificiaelll iingtence (AI), as an advanced data analysis i tsoo inl,tensively utilized to overcome the high complexity of datatte prans obtained. Recent advancements of usingic aiarltif intelligence in EoC prediction is reviewed by Raufeiq et al. [Rafique 2021]. Machine learning (ML) and ML-based classification are one of the importanbts seuts of AI based tools. The review paper froml P eatte al. [Patel 2020] is summarizing numerous challenges in omics dastsaist aed precision medicine and prediction of EoC. In general, however, the apptiloicna of AI has become possible in the field. Altered N-glycan profiles have been observed inio vuasr cancer types [Peric2022][Wang 2023][Thomas 202]1 and the aberrant profile is often considered ttoen ptoial diagnostic or therapeutic target. Capillary electrophoresis is an efficiental aytical method that enables the separation and identification of various molecules - in our ca Nse-g,lycan molecules [Lu 2018]. From practical point of view, the direct output of an analysis is the slole cda electropherogram, which is the acquired deotre scitgnal in the function of the time and usually containsns ceocutive peaks, each of them representing a single component (molecule) of the analyzed sample. Inre m foormal words, the electropherogram demonstrate the separation and quantification of the analyzeodle mcules. The evaluation and data interpretatio ann of electropherogram is a complex task usually doneu malalyn by experienced separation scient [isLtus 2018]. However, manual peak identification has some m darjaowr backs, especially in cases where, the strulctura or clinical information is encoded in a complexm foart i.e. not in the alteration of one or few indiuvail peaks but held in the change of the ratio of numuser poarallel peaks. Such alteration can be obse inrv tehde migration time of a molecule or in the peak areao.m Cputer assisted peak evaluation is just recently introduced for liquid chromatography metho [Sdsingh 2023][Risum 2019]. Furthermore, computer assisted or machine learning (ML) based data interpreta otifo cnapillary electrophoresis based analyses is in als itso emerging phas [eMészáros 2020]b. The paper by Iwamura et al. described a metho wd,h in ch immunoglobulin glycosylation togetherwith machine learning based statistics are utili fzoerd precision diagnosis of urological diseasems.il Sairly, Demirhan et al. demonstrated the high diagnostwice pro of the ML supported mass spectrometry based glycomics. However, none of the cited papers fodcu osne the prediction of the patient response for chemotherapy, which is a more challenging and ceomx pl roblem compared to the analysis and / or diagnosis of a certain tissue. Taniguchi, N. et al. Reviewed the role of N-glyca anss cancer biomarkers in progression and metastasis of tumours. However, apparently no work is focused yet on thoeten ptial to use glycan biomarkers in EoC prediction. Albeit advanced treatments, the overall incidenecve l l of lung cancer remains significa [Sntchabath2019]. THE DISCOVERY ACCORDING TO THE PRESENT INVENTION The present inventors have surprisingly found c thoamtplex albeit reliable correlations exist between serum N-glycome changes and the effectivenesse omf o chtherapy during lung cancer treatment and, by correlating the data with tumor sizes in latere staogf the disease, came to the unexpected conclu thsaiotnprediction of the effectiveness of chemotherapaytm treent in a patient with cancer disease is pos.s Tibhlemethod of the invention is applicable even with tohuet anamnesis of the patient involved. In the study of the inventors laser-induced fluocerensce detection equipped capillary electrophoresis analysis (CE-LIF) was used along with ML-based d partoacessing. The unique combination of the dataobtained by high-resolution N-glycan analysis and st a te-of-the-art data mining tool allowed theexploration of the relationship between clinicalra pmaeters and the alteration of the asparagined linke carbohydrates of the serum glycoproteins. Based on the revealed strong association betweee sntr tuhctural changes in the N-glycans and the outcomes of the chemotherapy treatments, the o obfje tchte present invention is to provide a methord fopredicting the effectiveness of chemotherapy treantm in lung cancer utilizing serum N-glycome aniasl.ysTo achieve the above-mentioned objective, the itnovresn have performed systematic experimentalwork, which has resulted in the present invent Tiohne. present invention is based on the finding t th earteis a relationship between the structural change thse in N-glycans and the outcomes of the chemotherapy treatments, more specifically, between the altoenra otif the asparagine linked carbohydrates of thruem se glycoproteins and the clinical parameters. The discriminating power of the method according th teo present invention was assessed in terms of the AUC values of the receiver operating charascttiecr (iROC) curves, and the results show unexpeyctedl great discrimination property (ROC > 0.9). There tbhyis, novel combined bioanalytical and AI method provides a precise and rapid tool for predictineg e thffectiveness of chemotherapy. BRIEF DESCRIPTION OF THE INVENTION The invention relates to a method for predictineg e thffectiveness of chemotherapy treatment in apatient with cancer disease, preferably lung ca dnicser ase, in particular as discloses hereinbelow.The invention also relates to a data carrier, prarebfley as a computer readable medium, more preferably a non-transitory computer readable mmed,iu storing machine executable instructions forexecution by a processor of the steps of the coimngpar nd optionally the assessing and / or analystienpg( s )of the method of the invention. In particular the invention relates to a method p forerdicting the effectiveness of chemotherapy treatment in a patient with cancer disease, prbelfyer luang cancer disease, said method comprising a.1. providing a first serum sample of the pat tieanketn at time one, a.2. providing a first serum sample of the pat tieanketn at time two, wherein said patient is treated by chemotherap thye in interval between time one and time two, b.1. obtaining a first N-glycome analysis resuoltm fr the first serum sample, to obtain a first glycan structure features pattern, b.2. obtaining a second N-glycome analysis resroumlt f the second serum sample, to obtain a second glycan structure features pattern, c. comparing the first glycan structure featurete prant and the second glycan structure feature pna ttoter assess changes in the patterns relating to starulct huarnges in the N-glycome of the patient, byy cinagrr out classification task to predict 'regression', 'perosgsrion' and 'stationary' states for characteri tzhieng respective chemotherapy outcome in the patient. In a particular embodiment the N-glycan profile co isrrelated with the predicted outcome of the treanttm of a cancer. The comparing and / or assessing, thereby the evinaglua st ep of the invention is useful and leads areliable result and thereby allows therapeutics dioenci even without any knowledge of patient data or anamnesis. In a preferred variant no further o (it.hee.r than obtained from the sample, preferablmy fr thoe methods defined in the Brief description part)e pnatti data or anamnesis are / is used in the metho thde of invention. In a particular embodiment step c. comprises coimngpa thre first glycan structure feature pattern t ahned second glycan structure feature pattern to assheasnsg ces in the patterns relating to structural cehsan ing the N-glycome of the patient, by carrying out threee ipnedndent binary classification task with classisfyer characterizing at least regression and progres osfio thne disease, preferably regression, progres asniodn stationer proceeding of the disease. Preferably in the comparing step the following s cilfaysers are applied 'regression', 'progression' and'stationary', relating to / characterizing the restipvec chemotherapy outcome in the patient.In a preferred embodiment the method comprises d. assessing the changes in the patterns rela ttheed f tiorst and second serum sample to assignl tehvea rnet classifier thereby predicting the effectiveness ch oefmotherapy treatment in said patient. In a preferred embodiment the method compriseson oapltliy the assessing step comprises analysing whether the changes in the patterns h aarera c terized by 'regression', 'progression' or'stationary' classifyers thereby predicting thec etfifveness of chemotherapy treatment in said pta.tien Thereby the result of the analysis is evaluated. In a preferred embodiment in the comparing stepo,n u cparrying out the classification task, the pastter are compared with patterns in a database of gl sytcrauncture feature or characteristics pattern. In a preferred embodiment the treatment of thee pnatt iis dependent on the result of the method forpredicting the effectiveness of chemotherapy treantm , i.e. on the N-glycome analysis results.2. In a preferred embodiment the N-glycome anal ryessisults are obtained by separation of N-glycans from the samples; and preferably glycan structueraetu fres are assigned to the separated N-glycans. peak Optionally high performance liquid chromatographHyP (LC) can be applied with detection of N-glycans. In an embodiment detection is possible by masstr sopmeectry (MS). In an embodiment structure features are assign tehde t soeparated N-glycan peaks. Preferably, the N-glycome analysis is carried oyu cta bpillary electorphoresis list of the relevanytc galns for the classification task are selected by peaeka. ar Preferably the N-glycans are separated by capil elalercytrophoresis measurement on the samples, and the glycan structure features / characteristicsh aere re tlative area of peaks of the separated anvda rnetle N- glycans. In an embodiment the structural features are tlhaetiv re area of peaks of the separated and rele Nv-antglycans. In a highly particular embodiment the relevant psea rke those having >1% peak area. In a particular embodiment the list of the relev galnytcans for the classification task are selecteitdh w peak area > 1%. Preferably, the N-glycan profile is correlated w tihthe outcome of the treatment of a cancer 3. In an embodiment binary classification is apdpl tioe predict the efficacy of chemotherapy treatm, ent wherein, based on the relative peak area datae of se thparated glycans. Binary classification taskes ar categorized as regression, progression, or staryti.ona Preferably, the capillary electrophoresis datai onbetda before and after the chemotherapy treatment arused to assess a chemotherapy treatment efficac pyat i ennts; preferably, the glycan structure featurepattern(s) is / are obtained as a system of data. P sreetfserably the glycan structure feature patte)r ins( / sare high resolution N-glycomics data. Preferably the classifier selection method is ste dlec from the group of ML methods consisting ofa Support Vector Machines (SVM) classifier to fi and hyperplane in a high-dimensional space; a Quadratic Discriminant Analysis (QDA) classifi teor provide flexibility in situations where classes have different variances; a Random Forest classifier to build multiple deocnis tirees during its training; an eXtreme Gradient Boosting classifier to assem wbelaek decision trees sequentially through boosting method, thereby providing a robust ensemble cliacsastiiofn model; a Neural Network machine learning model to cap ctuormeplex patterns and relationships in data. The skilled person will understand that any othpeprro apriate classifier selection method may be used, and tested by the AUC value, once an ROC curvet i usp s from the data obtained. 4. Preferably a feature selection task is appliheder weby the appropriate glycans useful in the diastgicno analysis are identified based on glycan structeuaretu fre pattern(s) as a system of data sets. Pbrelyfe thrae glycan structure feature pattern(s) is / are higshol ruetion N-glycomics data. Preferably the appropriate glycan peaks are sedle bcyte one or more method selected from the group consisting of a BayesSearch method to predict the performan dcieffe orfent hyperparameter configurations, preferably directing the search for glycan peaks towards psroinmgi regions in the hyperparameter space; a Sequential Forward Selection (SFS) method asat aur fee selection technique to improve model performance. The skilled person will understand that any appiraotepr feature selection technique may be used, and thereby the appropriate glycan peaks selected w chaicnh be correlated with or which has effect on the classifier from the set of glycan peaks. Preferably Quadratic Discriminant Analysis (QDA) a ipsplied. 5. In a preferred embodiment the list of the renletv galycans for the classification task, preferably binary classification task are selected. In a preferred embodiment the glycans are sele fcrtoemd the group consisting of the followingasparagine-linked glycan structures: G1: FA4BG4[3,3,3,3]S4, G2: A2G2[6]S2, G3: FA3G3[6]S3, G4: A2G2[3]S2, G5: A2BG2S2, M3, G6: FA2G2S2, G7: FA2BG2S2, G8: FA2[6] G1S1, G9: A3G3[3]S2, G10: A2G2[6]S1, G11: A2BG2S1, G12: FA2G2S1, G13: FA2BG2S1, M7, G14: A4G4[6]S2, G15: FA2, M6, G16: FA2B, G17: FA2[6]G1, M7, G18: FA2[3]G1, G19: FA2B[6]G1, M8, G20: FA2G2, G21: M9 or structures as disclosed in Table 1 4. Preferably, the list of the relevant glycans th foer binary classification tasks comprises at le 2,a 3st, or 4 glycans selected from the group consistin tghe of following glycans: G6, G12, G13, G21 preferably comprises at least 2, 3, 4 or 5 glyc saenlescted from the group consisting of the following glycans: G6, G12, G13, G20, G21. In an embodiment the following preferred combinnastio of glycans comprising at least the following glycan pairs are selected: G6, G12; G6, G13; G6, G20; G6, G21; G12; G13; G12, G20; G12, G21; G13, G20; G13, G21; and G20, G21; OR the following preferred combinations of glycans cporimsing at least the following glycan triplets are selected: G6, G12; plus one, two or three glycans selectoemd f Gr13, G20 and G21; G6, G13; plus one, two or three glycans selectoemd f Gr12, G20 and G21; G6, G20; plus one, two or three glycans selectoemd f Gr12, G13 and G21; G6, G21; plus one, two or three glycans selectoemd f Gr12, G13, and G20; G12; G13; plus one, two or three glycans selecrtoemd f G6, G20 and G21;G12, G20; plus one, two or three glycans selecrtoemd f G13, G6 and G21;G12, G21; plus one, two or three glycans selecrtoemd f G13, G20 and G6;G13, G20; plus one, two or three glycans selecrtoemd f G12, G6 and G21;G13, G21; plus one, two or three glycans selecrtoemd f G12, G20 and G6; andG20, G21; plus one, two or three glycans selecrtoemd f G13, G12 and G6.More preferably, the list of the relevant glycanosr p frogression comprises at least 2, 3, 4 or ast l 5eaglycans selected from the group consisting of o thlleow f ing glycans: G3, G4, G6, G8, G12, G1 G3,16, G20, G21, or preferably from G3, G4, G6, G8, G12, G1 G3,20, G21. Particularly preferably, the list of the relevanlytc gans for the binary classification tasks compsris aet least 2, 3, 4 or 5 glycans selected from the gr coounpsisting of the following glycans: for regression: G1, G2, G6, G12, G13, G14, G15, G17, G19, G20, G pr2e1f,erably G13, G14, G15, G17, G19, G20, G21, and / or for progression, G3, G4, G6, G8, G12, G13, G16, G20, G21, prefer farbolmy G3, G4, G6, G8, G12, G1 G3,20, G21, and / or for stationary status: G2, G3, G6, G11, G12, G13, G16, G18, G19, G20, G or21 p,referably from G3, G6, G11, G12, G13, G16, G18, G19, G21. In a particular embodiment the glycans are sele fcrotemd the or structures as disclosed in Table 5 5. In a preferred embodiment the N-glycan profileo isrre clated with the outcome of the treatment of a cancer. 6. In a preferred embodiment of the method define pdaragraphs 1 to 5, said classifier is trainetdh wi using independent data sets generated by randopmlitltyin sg the original data set for each classifiiocnat tasks. In a preferred embodiment splitting the originatla d saet into 80% training and 20% test data for each classification tasks. 7. In a preferred embodiment of the method defi ine pdaragraphs 1 to 6, wherein the N-glycome(pattern) is correlated with the outcome of theat tmreent of a cancer selected from the following gprou f cancers: preferably selected from the group consisting ofn- nsmoall cell lung cancers, preferably adenocarcinoma, squamous cell carcinoma, and c laerlgl cearcinoma. The skilled person will be aware of other lung cearnsc which can be treated or diagnosed by a method as disclosed herein. 8. Preferably,the trained classifier is testedn odnep i endent data sets generated by randomly sgpl tihttein original data set for each classification task. Preferably, a trained computer algorithm is userd cl faossification. In a particular embodiment, during classification co arrelation analysis of the relative percentageea ar of glycans is also performed. In a particular embodiment said classifier algomrit chomprises using machine learning or deep learning algorithms.9. Preferably, said relevant glycans may undergoi aodndailt feature selection, 10. In a preferred method of the invention the aredae urn curve (AUC) values of the receiver operating characteristic (ROC) curves provided by the claiesrsi afre at least 0,75, preferably at least 0,8,e morpreferably at least 0,85, particularly preferabtly le a st 0,9.11. In highly preferred embodiments Quadratic dimisicnrant analysis (QDA) is used for classification. 12. In a certain embodiment a glycan structuretu frea / characteristics pattern obtained from a heyalthhuman serum sample is also applied as a control. In a further embodiment the glycan structure feea / tcuhraracteristics pattern are obtained by capillary electrophoresis and a representative electrophraemrog from a control healthy human serum sample is applied as a control. 13. The invention also relates to a method ofm trea ntt wherein a method according to the inventiogn. e.according to any of paragraphs 1 to 12 is carruietd, a ond wherein chemotherapy outcome in the patient iss cifliaesd as 'regression', the chemotherapy treatment is coendti.nu In a preferred embodiment of the method the folnlogw siteps are carried out: a.1. taking a first serum sample of the patienetn ta akt time one, a.2. taking a first serum sample of the patienetn ta akt time two, wherein said patient is treated by chemotherap leya ast in the interval between time one and time, twob.1. obtaining a first N-glycome analysis resuoltm fr the first serum sample, to obtain a first glycan structure features pattern, b.2. obtaining a second N-glycome analysis resroumlt f the second serum sample, to obtain a second glycan structure features pattern, c. comparing the first glycan structure featurete prant and the second glycan structure feature pna ttoter assess changes in the patterns relating to starulct huarnges in the N-glycome of the patient, byy cinagrr out classification task to predict 'regression', 'perosgsrion' and 'stationary' states for characteri tzhieng respective chemotherapy outcome in the patient, preferably, the N-glycan profile is correlated w tihthe predicted outcome of the treatment of a ca,ncer wherein, if the chemotherapy outcome in the patient is cifliaesds as 'regression', the chemotherapy treatmsent i continued, if the chemotherapy outcome in the pnatt iise classified as 'progression', the chemotherapy treatment is shifted to chemotherapy treatmenta wi dthifferent active agent, and, if the chemotherapy outcome in the patient is cifliaesds as 'stationary’, the chemotherapy treatmesnt ieither continued or shifted to chemotherapy treantm weith a different active agent, optionally depienngdfrom a further measurement. Preferably a measurement according to any of an pya orafgraphs 1 to 12 is carried out as a further measurement in a further time point and the fut ihmeer point is considered as the second time-ponindt a either first or second time-point is considered th aes first time-point. Alternatively, an independent measurement of prtiendgic the chemotherapy outcome is carried out. 14. The invention also relates to a method ofm trea ntt wherein a method according to the inventiogn. e.according to paragraphs 1 to 13 is carried out, and wherein chemotherapy outcome in the patient iss cifliaesd as 'progression' the chemotherapy treatment is abaendd aon d / or shifted to another treatment e.g. anotherchemotherapy treatment. 15. The invention also relates to a method ofm trea ntt wherein a method according to the inventiogn. e.according to paragraphs 1 to 14 is carried out, and wherein chemotherapy outcome in the patient iss cifliaesd as 'stationary' the chemotherapy treatment is the chemotherapy treatment is continued, or the chemotherapy treatment is abandoned and / orerd al ttoe another treatment e.g. another chemotherapy treatment, and / or the method of the invention is repeated eitherr aft seecond chemotherapy treatment either by the sam chemotehapy or by an altered chemotherapy. 16. The invention also relates to a system comnpgrisi adata carrier, preferably as a computer readabeldeiu m , more preferably a non-transitory computerreadable medium, storing machine executable intsiotrnusc for execution by a processor of the step tshe of comparing and optionally the assessing and / or asinnagly step(s) of the method of the invention, or a computer comprising said data carrier and optioynall an apparatus for separating and / or analysing slyacidan gs from said samples. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1: CE-LIF separation of PNGase F released and APaTbSel-eld N-glycans from a healthy controlhuman serum sample. Glycans are numbered andf i edden utising the Oxford notatio [nHarvey 2011] fortraceability. Peaks: G1: FA4BG4[3,3,3,3]S4, G2: A22[6G]S2, G3: FA3G3[6]S3, G4: A2G2[3]S2, G5: A2BG2S2, M3, G6: FA2G2S2, G7: FA2BG2S2, G8: FA2 G

[0061] S1, G9: A3G3[3]S2, G10: A2G2[6]S1,G11: A2BG2S1, G12: FA2G2S1, G13: FA2BG2S1, M7, G A14 G: 4[6]S2, G15: FA2, M6, G16: FA2B,G17: FA2[6]G1, M7, G18: FA2[3]G1, G19: FA2B[6]G1,8 M, G20: FA2G2, G21: M9. Separation conditions: 40 cm effective capillary length (50 ctomtal length), 50 μm ID / 365 μm OD bare-fused silica capillary; Applied separation voltage: 30 kV (0. m17in ramp time) in reversed polarity mode. LIF detitoenc (excitation: 488 nm / emission: 520 nm); Separateiomnp terature 30°C. Injection: water pre-injection s 5.0 at 1.0 psi, followed by 2.0 kV / 2.0 s sample injoenct.iFigure 2: Pie chart demonstration of the dataset (i.e. c lla bsesls) distribution in the case of the threedifferent classification tasks. Approximately 33e%p r esent True labels, while around 66% represelsnet Falabels for all cases. Figure 3: Relative peak intensity distributions of N-glyc satnructures grouped by the classification tasks. Fig.4: ROC curves generated by the fine-tuned QDA alhgmor,it in case of prediction of A) regression, B) progression, and C) stationary. ABBREVIATIONS APTS: 8-aminopyrene-1,3,6-trisulfonic acid AUC: area under curve CBP: carboplatin CDDP: cisplatin CE-LIF: capillary Electrophoresis with Laser-Indudc Feluorescent Detection CRT: chemoradiotherapy ETO: etoposide GEM: gemcitabine PEM: pemetrexed QDA: quadratic discriminant analysis ROC: receiver operating characteristic SDS: sodium dodecyl sulfate SFS: sequential forward selection SNEC: small cell neuroendocrine carcinoma TAX: paclitaxel THF: tetrahydrofuran TXT: docetaxel DEFINITIONS The term “sample” is meant herein to refer to as staunbce comprising a mixture of compounds prepared separation, e.g. for analysis, e.g. byilla crayp electrophoresis. The sample may be derivreodm, f for example, a substance obtained from an environntamle source, e.g. bodily fluid or tissue from aj seucbt, i.e. taken from said bodily fluid or tissue of t shuebject and optionally processed to prepare folry asinsa. The sample as used herein preferably comprisesoh cyadrbrates, preferably glycans of biological or biotechnological interest, as defined, describe edx oermplified herein. A sample is typically a reaction mixture or a p tahretreof comprising said carbohydrates, preferably oligosaccharides and labeling agents. The labe alginegnts used herein are charged molecules, e.giv peolysit or negatively charged, wherein preferably the leadbe clarbohydrates, once reacted with the labelinegnt a,g become charged, e.g positively or negatively chda,r rgeespectively. “Carbohydrates” in a broader sense are compounvdinsg ha the stoichiometric formulanC(H2O)n, or derivatives, e.g. those having a major moiety (eprraebfly of at least 30%, 50% or 70% or 80% or themolecular weight of the whole compound) said ma mjor iety having the stoichiometric formulan( CH2O)n.“Carbohydrates” are preferably are aldoses or keest.o Tshe term carbohydrate includes monosaccharidesand oligosaccharides and polysaccharides as we sllu abs tances derived from monosaccharides byreduction of the carbonyl group (alditols), by oaxtiidon of one or more terminal groups to carbox aylcicids, or by replacement of one or more hydroxy group(ys) a b hydrogen atom or a substituent, in particular asubstituent having a molecular weigh of at mos at o mfonosacharide, eg. a functional group, in palartricuan amino group, thiol group or similar groups.ls Ito a includes derivatives of these compounds. In a sense, “glycans”, “oligosaccharides” and “psoalcycharides” are used herein interchangeably and mean a carbohydrate comprising or having multipoleno msaccharide units, preferably linked glycosidlyic.al In a narrower sense “glycans”, “oligosaccharidensd” a “polysaccharides” are as defined by IUPACCompendium of Chemical Terminolog [IyUPAC, (2019), #3]0 In a preferred embodiment “glycans” are glycan bsi o lfogical or biotechnological interest.In a broader sense glycans may comprise monosaidcecsh.ar In a preferred embodiment the term “glycan” relates to the carbohydrate portion ofly aco gconjugate, such as a glycoprotein, glycolip oird, a proteoglycan. Glycans can be homo- or heteropolsym ofe mr onosaccharide residues (may be composed of a single type of monosaccharide residue or fromtip mleul type of monosaccharide residue). Glycans b cean linear or branched. Glycans can be typically N-elidn-kglycans or O-linked-glycans or glycosaminoglysc.an “Detecting” as used herein is understood broasdl oyb ataining an observation regarding a substanceor compound of interest (preferably an analyte) a, a ressult of the separation method on a sample.A “detector” is a device for detecting and which lo icsated typically in connection with a specificte si of the CE capillary to detect compounds of inte mreisgtrating therein. “Binary classification” is a process used in maceh lienarning. The goal of the “binary classification” is to classify input data into one of two possi cbaletegories / classes. This process includes prnegdic ati binary outcome (yes / no, true / false) based on tphuet in attributes. It is used in a different appliocantsi (e.g. disease diagnosis). The term “comprises” or “comprising” or “including a”re to be construed here as having a non- exhaustive meaning and allow the addition or inevomlvent of further features or method steps or components to anything which comprises the liseteadtu fres or method steps or components. “Comprising” can be substituted by “including” if the practicfe a o given language variant so requires or canm biete ldi to “consisting essentially of” if other members or cpoomnents are not essential to reduce the invenotion t practice. The singular forms “a”, “an” and “the”, or at lea “sat”, “an”, include plural reference unless the context clearly dictates otherwise. DETAILED DESCRIPTION OF THE INVENTION The present invention relates to a method to ptre thdeic effectiveness of chemotherapy treatment in lung cancer by monitoring the serum N-glycome oef p thatients combined with data analysis. The study disclosed in the present description involvedy th-tihrtree lung cancer patients undergoing chemothyerap treatments. Serum samples were taken before aenrd th aeft treatment. The N-linked oligosaccharidese wer enzymatically released, fluorophore-labeled, sutbejdec to magnetic bead-mediated cleanup and analyzed by capillary electrophoresis with laser-inducedor fleuscent detection. The samples can be pretreated following a comprseivhen protocol as described [ inSching 2023]and [Risum 2019]. The resulting electropherograms were thoroughlyce psrosed and evaluated by classifiers, i.e., utilizing an algorithm to allocate the data intoo tw (binary) classes. The classifier analysis met rheovdealeda strong association between the structural cha ing thees N-glycans and the outcomes of the chemaopthyertreatments (ROC > 0.9). This novel combined bioyatnicaall and AI method provided a precise and rapid tool for predicting the effectiveness of chemothpeyr.a In a preferred embodiment artificial intelligencaes bed data analysis is used to analyse N-glycome data of the patient. The inventors made significant efforts to reveael c thorrelation between structural changes in theserum N-glycome and chemotherapy treatment resp ionn lsueng cancer patients, utilizing a classificantio model workflow in conjunction with CE-LIF analysi As. comprehensive analysis of 21 asparagine-linked glycan structures was conducted as shown in Ta.ble 1 Table 1 - Relevant asparagine-linked glycan struecstu for the binary classification tasks G1: FA4BG4[3,3,3,3]S4, G2: A2G2[6]S2, G3: FA3G3[36,]S G4: A2G2[3]S2, G5: A2BG2S2, M3, G6: FA2G2S2, G7: FA2BG2S2, G8: FA2[6] G1S1, G9: A33[3G]S2, G10: A2G2[6]S1, G11: A2BG2S1, G12: FA2G2S1, G13: FA2BG2S1, M7, G14: A4G4[6]S2,5 G: F1A2, M6, G16: FA2B, G17: FA2[6]G1, M7, G18: FA2[3]G1, G19: FA2B[6]G1, M8, G20: FA2G G2,21: M9. Identified N-glycan structures (nomenclature followed the gluinidees of the Consortium of Functional Glycomics, using the Oxford notation)a [rHvey 2011]. Entry numbers correspond to the peak numbers in the Figures. Symbols: Sialic acNid-a (cetylneuraminic acid); Galactose; N- acetylglucosamine; Mannose; Fucose. Peak No. Peak ID Structure Peak No. Peak ID Structure G18, FA2[3]G1 Peak No. Peak ID Structure Binary classification tasks were established toci sfipceally predict the efficacy of chemotherapy treatment, categorized as regression, progres osrio snta,tionary, based on the relative peak area o dfa thtae separated glycans. It was not at all foreseeable whether it is poses aibtl all to carry out this classification with alia reble result, i.e. to use the capillary electrophoresaitsa d obtained before and after the chemotheraptym treenat to assess a chemotherapy treatment efficacy in psa.tient Among the most often utilized ML methods differe cnlatssifiers play important role. The Support Vector Machines (SVM [)Hearst 199]8 classifier aims to find a hyperplane in a high- dimensional space that best separates data pnotiont dsif iferent classes. This hyperplane is positdione to maximize the margin, which is the distance betnw tehe hyperplane and the nearest data points from each class, known as support vectors. Them oaplti hyperplane provides a robust decision boundary, making SVM suitable for various classaitfiiocn tasks, especially when dealing with complex relationships in the data. SVM is effecti ivne handling both linear and non-linear relationships in the data, however, its basic fo isrm searching for linear separation of classes. The Quadratic Discriminant Analysis (QDA) classrif [ieHastie 2021][Balabin 2010] assumes a separate covariance matrix for each class, pirnogvi fdlexibility in situations where classes have different variances. This flexibility enables QDAo m t odel more complex decision boundaries,making it suitable for datasets with non-linearat rieolnships between features. QDA essentially calculates class-specific quadratic surfaces totin dgiusish between different classes, providing a robust method for classification in multivariateta dsaets with varying covariance structures among classes. The Random Forest classif[iBer eiman 2001] is a versatile machine-learning ensembletechnique. It builds multiple decision trees dur iitnsg training. Each tree is constructed indepenlyd,ent introducing randomness in the training processs binyg u subsets of data and features. Random Forest is adept at handling complex relationships, mininmgiz overfitting, and providing high predictive accuracy. The eXtreme Gradient Boosting classi[fCiehren 2016][Cao 2023] assembles weak decision trees sequentially through boosting method, pronvgid ai robust ensemble classification model. Each tree aims to minimize the error of the loss funnct oiof the previously built tree. XGBoost is designed to optimize performance and minimize overfittingro,v piding high efficiency across various tasks. A Neural Network [Müller 1995] is a machine learning model inspired by the sutrruect and function of the human brain. It consists of intenrnceocted neurons, organized in layers and functions as a "black box" in the sense that the internalk winogrs and transformations of the data within the network can be complex and not easily interpreta bbyle humans. Neural networks can capture complex patterns and relationships in data, ma tkhinegm suitable for deep learning, however, their training generally requires a huge amount of data. The present inventors have investigated the apbpilliitcya of various classifiers in predicting changes in lung cancer stages based on N-glycan profiling. Among the fine-tuned classifiers, the Quadraticr dimisinant analysis (QDA) classifier proved to be the most effective method resulting in AUC valuefs 0 o.8918, 0.9039, and 0.9052 for regression, progression, and stationer stage changes, resepleyc.ti Wv hile the ensemble classifiers also generated satisfactory results, their accuracy slightly f sehllort of expectations. Thus, QDA was selected for constructing the finaold mel that underwent fine-tuning to improve the classification accuracy. To arrive at a preferred embodiment of the methoed p tresent inventors identified relevant glycan peaks for predicting treatment efficiency acrosls th arlee binary classification tasks. This is a u fereat selection task wherein the appropriate glycansu uls inef the diagnostic analysis, if any, has to buen fdo in a highly complex system of data sets. The present inventors applied multiple tools toec ste tlhe glycan peaks appropriate for the prediction analysis. BayesSearch method utilizes probabilistic models p troedict the performance of different hyperparameter configurations, directing the sea torcwhards promising regions in the hyperparameter space. The Sequential Forward Selection (SFS) method fe isat aure selection technique that improves modelperformance by iteratively adding one feature t aimt ae to the subset of selected features. It be wgin ths anempty set and evaluates the impact of adding eeaacthur fe on the model's performance. In this study, the inventors adapted and testede nrouums classifier methods as data evaluation tools for monitoring lung cancer patients’ response toem cohtherapy using high resolution N-glycomics data obtained by capillary electrophoresis coupled w laitsher-induced fluorescence detection. An initial result is described in Example 8. The combination of the SFS and brute force procee,d ausr described in Example 9, identified more relevant glycan peaks for predicting treatmentc ei,0ff0i ency across all three binary classificatiosnks ta. As a working hypothesis, initially the inventorsti acinpated that all three chemotherapy responsetypes can be accurately predicted using the samtae s deat. Thus, the performance of the QDA classti oficnawas first evaluated by using only those glycanc stutrrues, which were common in all three cases: p Gea6k,s G12, G13, G20, and G21. In this instance, the prmerafonce of the fine-tuned classification model woats n acceptably appropriate as it resulted the averargeea A Under Curve (AUC) of 0.7632. This inferiorperformance of the classifier suggested that thCe E proediction should be handled separately for eachresponse types utilizing the reduced datasetsd perodv biy the hybrid SFS and brute force algorithmu.s T,h the QDA classifier was employed and evaluated oen re thduced datasets for each binary classificaatisokns t. The discriminating power of the model was asse isnse tedrms of the AUC values of the receiver opergatin characteristic (ROC) curves provided by the QDAss cilfaier. The specific ROC curves for each classification task are plotted in Figure 4. The resulting AUC values exceeded 0,8, and in i cner etambodiments preferably 0,85, highly prefereably 0.9, thus, the algorithm successfulrlyed picted the effectiveness of the corresponding chemotherapy treatment, based on changes in thlyec Nan-g profile of serum samples taken before anedr aft therapy. This synergistic approach not only demroantesdt promising results in predicting chemotherapy efficiency through N-glycan analysis but also hightled the transformative power of AI in biological sample analysis. Classifier in machine learning is an algorithm th soartts data into one or more of a set of classes[Sarker 202]1. The classification task, which aimed to explohre t correlation between chemotherapyoutcomes and structural changes in the N-glycomofeile psr, was transformed into three independentr byina classification tasks to predict the effectivenesfs ch oemotherapy, with the following class labels of 'regression,' 'progression,' and 'stationary.' To identify the most suitable classification met,ho thde performance of 27 different classification algorithms was probed using their default paramse.t Qeuradratic Discriminant Analysis (QDA) showed thebest classification performance and thus it wasec ste dl to construct the final mod [eTlharwat 2016].Furthermore, the parametrization of the QDA alghomrit was fine-tuned for optimal classification accuyr.a Additionally, a combination of the Sequential Fereat Suelection (SFS) procedure [[2S6c]hüppstuhl 2019] and the brute force method were employed to idfyen thtie relevant N-glycan peaks with structural changes most effectively correlating the chemotphyeraesponse. Due to the limited number of records in the origl i dna taset, the fine-tuned QDA classifier was run100 times for the final evaluation. This involveadnd romly partitioning the entire original datasetto in separate training and test sets, and the resu plteinrfgormance metrics were calculated by averagineg thresults of the test sets. All in-house developetad a danalysis code was implemented in Python usinpgyte JurNotebook v7.1.1 [Toomey 2017]. EXAMPLES Materials used in Examples: Chemicals and reagents Sodium dodecyl sulfate (SDS) and Nonidet P-40 w freorme VWR (Radnor, PA, USA). Acetonitrile,glycerol, dithiothreitol (DTT), tetrahydrofuran (TFH), sodium cyanoborohydride (1 M in THF), and acceti acid were from Sigma Aldrich (St. Louis, MO, USA T)h.e Fast Glycan Labeling and Analysis Kit was fromBioscience Kft (Budapest, Hungary). The endoglydcaos ei PNGaseF for N-glycan release was made in-house as described [ iKnovács 2022]. Example 1 - General method Capillary Gel Electrophoresis A PA800 Plus Pharmaceutical Analysis System wieth 3 t2hKarat (version 10.1) data collection and processing software package (Beckman Coulter) wppalsie ad for the analysis of the releas Ne-dlinked APTS labeled glycan structures in CE-LIF mode usingm 40 e cffective length (50 cm total length), 50 µm I6D5 / 3 µm OD bare fused silica capillaries filled with HNRC-HO separation gel buffer (Bioscience Kft). The separations were accomplished by applying 30 kVctr eicle potential in reversed polarity mode (catho adte the injection side, anode at the detection side 3)0 a°Ct capillary temperature. A water plug pre-intijoenc (1.0 psi for 5.0 s) preceded the sample injectiyon ap bplying 2.0 kV for 2.0 s. Relative percentagea ar values of the separated peaks were calculatede b Pye thak Fit v4.12 Software (SeaSolve Software S Inacn., Jose, CA). Example 2 - Specimen collection Pathological samples were collected from lung cran pcaetients undergoing chemotherapy at the Department of Pulmonology in Semmelweis Hospitalis (kMolc, Hungary), following the appropriate ethical permissions (approval number: 23580-1 / 2E0K15U / (0180 / 15)) and with informed patient consents. Thirty-three patients of Caucasian descent withg l cuanncer, receiving diverse doses of chemotheriacpeut agents (see detailed information in Table 2), w inecreluded in the study. Serum specimens were obdtaineboth before and after each treatment session abnsdeq su ently stored at -80°C until processing.Table 2: Lung cancer patients undergoing chemotherapy Patient Age Sex Histology Stage Applied Chemotherapy 1 59 male squamous cell carcinoma IV first-linelia ptaivle TAX-CBP 2 55 male squamous cell carcinoma IV first-linelia pativle GEM-CBP 3 67 male SNEC IIIB first-line palliative CBP-ETO 4 58 female adenocarcinoma IV first-line palliat PivEeM-CDDP 5 70 male adenocarcinoma IV first-line palliativEeM P-CDDP first-line palliative Bevacizumab 6 69 male adenocarcinoma IV + CBP + TAX 7 61 female adenocarcinoma IV first-line palliat GiveEM-CBP 8 47 male adenocarcinoma IIIB first-line palliati GveEM-CBP 9 68 male adenocarcinoma IV first-line palliativeEM G-CBP 10 53 male SNEC IV first-line palliative CBP-ETO 11 69 male adenocarcinoma IA adjuvant GEM-CBP adenosquamosus 12 58 male carcinoma IIIB first-line palliative GEM-CBP 13 62 male SNEC IV first-line palliative CDDP-ETO 14 61 male squamous cell carcinoma IIIB first-l pinaelliative GEM-CBP 15 69 female SNEC IIIA first-line palliative CPBT -EO 16 69 female adenocarcinoma IIIA adjuvant GEM+CBP 17 74 male adenocarcinoma IIIA adjuvant GEM+CBP 18 66 female adenocarcinoma IV first-line palliaeti GvEM-CBP 19 63 male adenocarcinoma IV first-line palliat GiveEM-CBP 20 67 male squamous cell carcinoma IIIA first-l pinaelliative CBP-TXT 21 47 male adenocarcinoma IB adjuvant GEM+CBP 22 62 male adenocarcinoma IIIB first-line palliaeti GvEM-CBP 23 62 male SNEC IV first-line palliative CBP-ETO 24 61 male adenocarcinoma IB adjuvant GEM+CBP 25 72 male squamous cell carcinoma IV first-linellia ptaive GEM-CBP 26 56 male squamous cell carcinoma IIIB first-l pinaelliative TAX+CBP 27 65 female adenocarcinoma IA first-line palliaeti GvEM-CBP 28 70 male squamous cell carcinoma IV first-linellia ptaive TAX-CBP 29 63 female SNEC IIIA first-line palliative CBP-EOT 30 73 male adenocarcinoma IIIA first-line palliaeti GvEM-CBP 31 72 female adenocarcinoma IV first-line palliaeti PvEM-CBP 32 59 female adenocarcinoma IIIA CDDP / TXT-CRT 33 65 male squamous cell carcinoma IIIB first-l pinaelliative GEM-CBP Age average: 63.4, Age median: 63, Age range: 47–74 Example 3 - Sample preparation The sample preparation protocol included denatounra,t Ni-glycan release, fluorophore labeling, and magnetic bead-mediated cleanup. Serum samples were diluted a hundredfold with HPgLraCd-e water and then denatured at 70°C for10 minutes by adding 2.0 µL of denaturation sonlut firo m the Fast Glycan kit (Bioscience Kft). Glycanrelease was achieved by adding 1.0 µL of PNGaszeyFm een (200 mU) to the reaction mixture followed by incubation at 37°C for 2 hours to ensure completgely dcosylation. The endoglycosidase digestion rioenact was stopped by adding the labeling solution, wh ciochntained 1.0 µL of 40 mM 8-aminopyrene-1,3,6- trisulfonic acid (APTS) in HPLC-grade water, 2.0 µofL NaBH3CN (1 M in THF), 10 µL of 50% acetic acid, and 8.0 µL of THF. The reaction mixture was incubated in a heatingck bl oovernight at 37°C in an open vial (all liquid evaporated) [Reider 2018], purified by using a magnetic bead-based appro [Vacáhradi 2014] and then analyzed by CE-LIF. All measurements were maderip inlic tates. Example 4 - Dataset Creation The labeled samples were analyzed by capillarytr eolpehcoresis with laser-induced fluorescence detection (CE-LIF) by using similar parametersn as an i earlier publication [Mészáros 2020]a[Mészáros 2020b]. The relative peak profiles from each serum sa,m appleplied in triplicates, were averaged before proceeding with the data analysis. The available sample size for the analysis waste lidm,i consisting of 98 samples from 33 lung cancerpatients. Due to the small sample size, an intnrigu rei search question was whether this sample saizse wsufficient for uncovering potential correlationsh.e T preprocessed dataset included the relative peak intensities of samples (21 attributes), the anonzyemdi patient identification code, and informationou atb the stage change. Based on the lung cancer stage ch aattnrigbeute, the following class labels were defi:ned "regression", "progression", and "stationer". Thlaess cification task was converted into three binary classification tasks, each one predicting a spcec hifiange of stage. Example 5 - Classification Methods and Fine-tuning In this study, five classification models were dleovpe d and applied: Support Vector Machine(SVM), Quadratic Discriminant Analysis (QDA), Ranmdo Forest (RF), eXtreme Gradient Boosting (XGBoost), and Neural Network (NN). The Support Vtoerc Machine was tested as a linear separator, and no kernel was applied during its execution. To increase the prediction accuracy, hyperparam oeptteimr ization was performed for each classifier. First, we determined the search space of the hyapraemrpeters to be tuned, and then Bayesian Optimonizati (BayesSearch) method was applied to fine-tune thoede mls in each classification case (regression, progression, stationer) separately. From a datlays aisna perspective, considering the limited amoufn dta ota,the BayesSearch method was executed 10 timese foryr e cvlassification case, randomly splitting their entdataset each time into disjoint training and teestst. s Through this repetition, the BayesSearch mdetho provided a more precise solution. Example 6 – Feature Selection and Handling Imbalance As each stages, “regression”, “progression” anadti “osntary” is probably indicated by different peaks or combinations of peaks, the SFS method was pmeerfdor separately for each classification task. The method was executed 50 times for each classifinca ttaiosk to identify the relevant attributes (N-glnyca structures). As a result of the SFS process, netawse dtas were generated for each classificationT tahseks.e new datasets included only the relevant attribu foters each case, aiming to increase the classifincatio accuracy of the models by fine-tuning the investetigda machine learning algorithms again. The distribution of the class labels is imbalanc ined the available dataset. After converting the problem into binary classification tasks, with aopxpimr ately 33% representing True labels and arou6n%d 6 representing False labels for all cases (regre:s Tsrioune labels: 35, False labels: 63, progressiorune: T labels: 30, False labels: 68, stationer: True labels: 3a3ls,e F labels: 65). Due to the limited amount of d aantda the unbalanced dataset, classifiers would have faceed ch thallenge of not being able to learn the mino crliatyss label efficiently. To avoid this problem the RandOovmerSampler was used for all classifier method tsh oentraining datasets. The ratio of the train-testt sp wlias 80% to 20%. We chose to use random overlisnagmpbecause, due to the small sample size, we wante advo tiod generating new synthetic data, which could potentially introduce incorrect samples and noise in aput to the classification algorithms. The RandomOverSampler balances class labels by rando smelelycting rows from the smaller class and duplicating them. Example 7 - Evaluation of the Resulting Classification Mosdel To evaluate the performance of the classificationde mls, the following quality metrics were applied: accuracy, sensitivity, specificity, F1-score, anUdC A value. The “accuracy” of the prediction measures the r oafti coorrectly classified samples to the total nurmbe of samples, indicating the overall correctnesshe of m todel. “Sensitivity” is the proportion of trueos pitives results in the positive class compared to the sfu tmrue o positives and false negatives predictions h.ig Aher sensitivity value shows better performance in dteintgec positive cases. “Specificity” measures the mel'osd ability to correctly identify negative instancesin ugs the values of true negatives and false posi.tiv Tehe “F1-score” combines the accuracy of positive prteiodnics as precision and sensitivity. It is particrulyla useful when there is an imbalance between the nru omfb peositive and negative instances. The AUC (Area Under the Curve) value represents the area unede Rre thceiver Operating Characteristic (ROC) curveic,h wh represents the “sensitivity” (true positive rateg)ai anst the false positive rate (1-“specificity”)h.e T higher AUC values indicate a better separation of posi atinvde negative cases, where the value of 1.0 renptrse aseperfect classification model. The AUC value is fure nqtly used in medical science, especially in theevaluation of diagnostic tests. Example 8 The SFS feature selection method was applied i ivter lyat on the original dataset for eachclassification task. The resulting relevant N-glnyc paeaks are presented in Table 3. As it is showarnio,u vsN-glycan structures are relevant in each classti ofinca task. Subsequently, new datasets were cre bat seeddon the resulting attributes for each classificati aosnk, and classifiers were executed on the coorrnedsinpg datasets. The results confirmed that the accurfac pyre odictions was adequate even if it was execuotned reduced datasets. Table 3. The relevant N-glycan structures for each clacsastiifoin task resulted from the SFS method. Classification task Peak IDs resulted from SFS Regression 3, 5, 6, 12, 13, 14, 15, 17, 19, 20, 21 Progression 3, 4, 6, 8, 12, 13, 15, 20, 21 Stationer 3, 4, 6, 11, 12, 13, 14, 16, 17, 18, 219,After applying the SFS method, the classifier ailtghomrs were fine-tuned again on the reduced datasets to further increase accuracy. Table 4en ptrses a comprehensive overview of the results o ffin aell- tuned classifiers for each classification task.a Psele note that results based on balanced datasreets we significantly more accurate than predictions getneedra without oversampling. Table 4. The results of fine-tuned classifiers for eachss cilfaication task using RandomOversSampler balancing technique. Classifier Metric Regression Progression Stationer SVM Accuracy: 0.6183 0.7437 0.6694 AUC: 0.6183 0.7437 0.6694 F1: 0.6429 0.7539 0.7025 Sensitivity (recall): 0.6977 0.7611 0.7908 Specificity: 0.6183 0.7437 0.6694 QDA Accuracy: 0.8175 0.8312 0.8367 AUC: 0.8918 0.9039 0.9052 F1: 0.8171 0.8311 0.8376 Sensitivity (recall): 0.8250 0.8436 0.8508 Specificity: 0.8918 0.9039 0.9052 RandomForest Accuracy: 0.7369 0.8468 0.7758 AUC: 0.7369 0.8468 0.7758 F1: 0.7459 0.8552 0.7868 Sensitivity (recall): 0.7800 0.9029 0.8338 Specificity: 0.7369 0.8468 0.7758 XGBoost Accuracy: 0.7494 0.8466 0.7927 AUC: 0.7494 0.8466 0.7927 F1: 0.7577 0.8549 0.8005 Sensitivity (recall): 0.7931 0.9046 0.8327 Specificity: 0.7494 0.8466 0.7927 Neural Network Accuracy: 0.7350 0.7750 0.7404 AUC: 0.7350 0.7750 0.7404 F1: 0.7374 0.7686 0.7356 Sensitivity (recall): 0.7554 0.7607 0.7358 Specificity: 0.7350 0.7750 0.7404 As it is shown in Table 4, the QDA classifier demstoranted high values on all metrics across all classes, indicating its effectiveness in predict hineg stage changes of lung cancer. All metrics th feor QDA classifier exceeded the value of 0.8, and the AUalCue vs are near or exceeded 0.9, showing a promorise f further research opportunities. The fine-tuned Q aDlgAorithm used the regression parameter of 0.0r01 fo each classification task. XGBoost and Random Forest classifiers providedr lo awcecuracy and AUC values. Although they achieved metrics around 0.84 in the classifica otifo pnrogression, their performance was slightly weera ink the other two classification tasks. SVM and ther Nael Nuetwork classifier exhibited the weakest mestr fiocr this dataset. As neural networks generally req aui lraerge amount of data for efficient learning, m thoedest performance was probably caused by the limited anmt o fu training data. While it is plausible that the other methods alsaon c provide a satisfactory result, in particular iftrained on a larger data set, the QDA classifietrh mode provided a surprisingly better performancen tha other methods. Example 9 In this example, the N-glycome of 98 serum samp (3le3s lung cancer patients, multiple samples collected from each patient depending upon therisro pneal therapeutic needs, see Example 2) werez aendaly using CE-LIF method to explore any structural cheasn ign their carbohydrate profile during chemothyerap treatments. Serum samples were collected after each treatmenstsio sn, and the asparagine-linked oligosaccharides were enzymatically released,e ladb welith a fluorophore (APTS) for downstream anasl.ysi Arepresentative electropherogram from a controal t he y human serum sample is shown in Figure1. The structural identification utilized directn minig of the GU database entries available in theca GlU v1.1c application linked with the GlycoStore daotalle cction [Jarvas 201]5[Jarvas 202]0. In this experimentonly glycans with greater than 1% relative peaka a wre re selected for downstream data analysis, mngeani21 peaks in our particular case. In this regard in thve ntors the selection and numbering of the gnlycastructures, i.e. peaks in the electropherogram w ceare ried out based on methods disclosed in earlierreports[Mészáros 2020]a[Mészáros 2020]b. The relative percentage area of the separatede alenvda rnt (>1%) peaks was calculated from the electropherograms of the serum samples collectefdore be and after the chemotherapy treatments. Subsequently, a QDA classifier was employed toy aznea tlhe correlation between changes in the relative peak areas, i.e. changes in the N-glycan profidle th aen effectiveness of chemotherapy treatment coaritzeegd as regression, progression, or stationary. Before classification, the unprocessed input deata w sas analyzed to explore its consistency. Theheterogeneous distribution of the datasets, i.uee. T argainst False of the three different binaaryss cilficationtasks is depicted in Figure 2. Please note, thme t “ebrinary” relates to the type of the classifie.re,., i todistinguish between positive and negative caseys, o wnhlile the three different tasks are the preodnict oifregression, progression, and stationary, respelyc.tive The descriptive power of the N-glycan profile cheasng was also evaluated using ANOVA test to shed light on the role of individual peak intensities t ihne classification tasks. Significant differenc weesre observed for peak IDs G6 and G15 only. In the c oafs Ge6, the disparity between measurements assodciate with progression and stationary class labels redsu ilnt a p-value of 0.0221, while for G15, the pu-veal between measurements associated with regression pr aongdression was p=0.0091. Nevertheless, the average probability value calculated using the ANAO tVest was notably high with the mean p-value of 0.5003, with a standard deviation of 0.2652. Thleati rvee peak intensity distributions of the N-glycan structures grouped by the classification tasks vi asruealized in Figure 3. It can be concluded from the thorough analysisig oufr Fe 3, that there is no significant difference in the distribution of the relative peak intensitines n ieither group i.e. in regression, progression s,ta otrionary. This is further confirmed by the ANOVA test, seeov aeb. Furthermore, correlation analysis of the relativeerc pentage area of glycans was performed as well. High correlations were found between the follow Nin-gglycan structure pairs: peaks G15 / G17 (r = 0.77), peaks G17 / G18 (r = 0.91), peaks G17 / G20 (r = 0. a8n7d), peaks G18 / G20 (r = 0.77). To identify the most relevant glycan structure cgheasn related to chemotherapy response in the binaryclassification tasks, SFS feature selection met whoads performed. The SFS algorithm was executedseparately for each classification task i.e.,h foer p trediction of progression, stationary or regiorens tsype of chemotherapy response. We refined the resultse o SfF thS method manually using the brute force method.This involved adding and removing glycan structu fre osm the selected set to identify the crucial galnycstructures for classification tasks. The hybrid S aFnSd brute force algorithm successfully removed the insignificant glycan structures, which did not croibnutte significantly to the chemotherapy respon frsoem,each dataset, resulting in a reduced datasette ads i lnis Table 5.Table 5: The list of the relevant glycan structures foer b thinary classification tasksClassification task Relevant glycan structure peaks Regression only by SFS G2, G3, G12, G13, G14, G15, G17, G1189,, G 20, G21 by SFS + brute force G1, G2, G6, G12, G13, G14,, G1157, G19, G20, G21 removed by brute force G3, G18 added by brute force G1, G6 Progression only by SFS G3, G4, G6, G8, G12, G13, G15, G16,, G1270, G21 by SFS + brute force G3, G4, G6, G8, G12, G13, G1260,, G21 removed by brute force G15, G17 added by brute force - Stationary only by SFS G3, G6, G8, G11, G12, G13, G16, G189,, G121 by SFS + brute force G2, G3, G6, G11, G12, G13,, G1168, G19, G20, G21 removed by brute force G8 added by brute force G2, G20 As a working hypothesis, we anticipated that arlele th chemotherapy response types can be accurately predicted using the same data set. Thus, the pmearfnocre of the QDA classification was first evalua bteyd using only those glycan structures, which were comnm in all three cases: peaks G6, G12, G13, G20, and G21. In this instance, the performance of thet fuinnee-d classification model was not acceptably apprpiarote as it resulted the average Area Under Curve (AUfC 0).7 o632. This inferior performance of the clasesrifi suggested that the EoC prediction should be han sdelpeadrately for each response types utilizinge thdeuc red datasets provided by the hybrid SFS and brute f aolrgcoerithm. Thus, the QDA classifier was employend aevaluated on the reduced datasets for each binla srysif cication tasks. The discriminating power oef thmodel was assessed in terms of the AUC valuese of re thceiver operating characteristic (ROC) curves provided by the QDA classifier. The specific ROCrv ceus for each classification task are plotted ignu Frei 4. The fine-tuned QDA classifier was executed on 1 r0a0n0domly generated, reduced datasets derived from the original set, to avoid overtraining or cvoenrgence to local minima, thus, to provide relia rbelseults.In general, the classifier was first trained, th te snted using independent data sets generated d boym ralynsplitting the original data set (80% training, 20 te%st) for each classification tasks. The result AinUgC values of the averaged 1000 runs were 0.8290 fgorres resion, 0.8295 for progression, and 0.8410 for stationary. ROC AUC analysis is a frequently useecdhn tique for analyzing the accuracy of diagnosetsicts t. The ROC curve is the plot of the series of trueit pivoes points (sensitivity) against the false povseiti points (1-specificity). An ideal ROC curve raise immedliyat teowards the upper left corner of the graph inadtiincg good (AUC > 0.8) or great (AUC > 0.9) discriminanti poroperty. In other words, higher AUC value of the ROC curve suggests greater discriminative power. INDUSTRIAL APPLICABILITY Based on the foregoing findings above, the pre isnevnetntion provides a method for predicting the effectiveness of chemotherapy treatment in lungce cra bnased on serum N-glycome analysis. The advantage of the method is that it enablesa arlny d eetection whether the chemotherapy startedin a patient diagnosed with lung cancer is reaflfleyc etive. Early detection of ineffective therapy p irsimarily beneficial for the patient, as it is possible toan cghe to another therapy in time. Furthermore, iatls iso advantageous from an economic point of view if weceo rgnize in time that an expensive therapy may be ineffective. REFERENCES Alghamdi HI, Alshehri AF, Farhat GN. An overviewf m oortality & predictors of small-cell and non-small cell lung cancer among Saudi patient Esp.i Jdemiol Glob Health. 2018 Mar;7 Suppl 1(Suppl 1):S1-S6. Bade, B.C. and C.S. Dela Cru Lzu,ng Cancer 2020: Epidemiology, Etiology, and Prevention. Clin Chest Med, 2020.41(1): p.1-24. Balabin, R.M., et al., Gasoline classification using near infrared (NIR) spectroscopy data: comparison of multivariate techniques. Analytica Chimica Acta, 671(1-2) (2010), 27–35. Bogos, K., et al., Revising Incidence and Mortality of Lung Cancer in Central Europe: An Epidemiology Review From Hungary. (2234-943X (Print)). Breiman L., Random Forests, Machine learning 45 (2001), 5–32. Cao, L., et al., Identification of Co-diagnostic Genes for Heart Failure and Hepatocellular Carcinoma Through WGCNA and Machine Learning Algorithms. Scientific Reports, 13(1) (2023), 14794. Chen, T., and Guestrin, C X.g,boost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International conference on knowledge discovery and data mining, (2016), 785– 794. Demirhan D.B. et al. Prediction of gastric cancer by machine learning integrated with mass spectrometry-based N-glycomics. Analyst. 2023 May 2;148(9):2073-2080. doi: 10.1039 / d2an02057b. PMID: 37009642. Harðardottir, H., et al., [Advances in lung cancer diagnosis and treatment - a review]. Laeknabladid, 2022.108(1): p. 17-29. Harvey, D.J., et al. S,ymbol nomenclature for representing glycan structures: Extension to cover different carbohydrate types. Proteomics, 2011.11(22): p. 4291-5. Hastie, R., et al. T,he Elements of Statistical Learning. Springer-Verlag, New York, 2001. Hearst, M.A., et al. S,upport vector machines. Intelligent Systems and their Applications, IEEE, 13(4) (1998), 18–28. Hirsch F.R., et al. L,ung cancer: current therapies and new targeted treatments. Lancet.2017 Jan 21;389 (10066):299-311. Iwamura H et al. Machine learning diagnosis by immunoglobulin N-glycan signatures for precision diagnosis of urological diseases. Cancer Sci. 2022 Jul;113(7):2434-2445. Jarvas, G., et al. E,xpanding the capillary electrophoresis-based glucose unit database of the GUcal app. Glycobiology, 2020. 30(6): p.362-364. Jarvas, G., M. Szigeti, and A. Guttma Gn,Ucal: An integrated application for capillary electrophoresis based glycan analysis. Electrophoresis, 2015. 36(24): p.3094-6. Kovács, N., et al. E, nhanced Recombinant Protein Production of Soluble, Highly Active and Immobilizable PNGase F. Mol Biotechnol. 2022 Aug;64(8):914-918. Lu, G., et al., Capillary Electrophoresis Separations of Glycans, Chem Rev 118(17) (2018), 7867- 7885. Mészáros, B., et al C.,omparative analysis of the human serum N-glycome in lung cancer, COPD and their comorbidity using capillary electrophoresis. Journal of Chromatography B, Volume 1137, 2020, 121913, Mészáros, B., et al M.,achine Learning Based Analysis of Human Serum N-glycome Alterations to Follow up Lung Tumor Surgery, Cancers 12(12) (2020), 3700. Müller, B., et al., Neural networks: an introduction. (1995) Springer Science & Business Media. Nagasaka, M. and S.M. Gadge Reol,le of chemotherapy and targeted therapy in early-stage non- small cell lung cancer. Expert Rev Anticancer Ther.2018 Jan;18(1):63-70. Nooreldeen, R. and H. Bac Chu,rrent and Future Development in Lung Cancer Diagnosis. Int J Mol Sci, 2021. 22(16). Patel, S.K., B. George, and V. R Aari,tificial Intelligence to Decode Cancer Mechanism: Beyond Patient Stratification for Precision Oncology. Front Pharmacol, 2020. 11: p.1177. Peric, L., et al. G, lycosylation Alterations in Cancer Cells, Prognostic Value of Glycan Biomarkers and Their Potential as Novel Therapeutic Targets in Breast Cancer. Biomedicines, 10(12) (2022), 3265. Rafique, R., S.M.R. Islam, and J.U. Ka Mzia,chine learning in the prediction of cancer therapy. Comput Struct Biotechnol J, 2021. 19: p.4003-4017. Reider, B., M. Szigeti, and A. Guttma Env,aporative fluorophore labeling of carbohydrates via reductive amination. Talanta, 2018. 185: p.365-369. Risum, A.B. and R. Bro U,sing deep learning to evaluate peaks in chromatographic data, Talanta 204 (2019), 255–260. Rosell, R., et al. P,redicting the outcome of chemotherapy for lung cancer. Curr Opin Pharmacol, 2006. 6(4): p. 323-31. Sarker, I.H., Machine Learning: Algorithms, Real-World Applications and Research Directions. SN Computer Science, 2021. 2(3): p.160. Schabath, M.B. and M.L. Cot Ce,ancer Progress and Priorities: Lung Cancer. Cancer Epidemiol Biomarkers Prev, 2019.28(10): p.1563-1579. Schüppstuhl, T., K. Tracht, and J. Roßman Tnag,ungsband des 4. Kongresses Montage Handhabung Industrieroboter. 2019: Springer Berlin Heidelberg. Singh, Y.R., et al. C,urrent trends in chromatographic prediction using artificial intelligence and machine learning, Analytical Methods 15(23) (2023), 2785–2797. Taniguchi, N. and Y. Kizuka C,hapter Two - Glycans and Cancer: Role of N-Glycans in Cancer Biomarker, Progression and Metastasis, and Therapeutics, in Advances in Cancer Research, R.R. Drake and L.E. Ball, Editors. 2015, Academic Pre ps.s 1.1-51. Tharwat, A., Linear vs. quadratic discriminant analysis classifier: a tutorial. International Journal of Applied Pattern Recognition, 2016.3(2): p.114850-. Thomas, D., A.K. Rathinavel, and P. Radhakrishn Aaltner,ed glycosylation in cancer: A promising target for biomarkers and therapeutics, Biochim Biophys Acta Rev Cancer 1875(1) (2021), 188464. Toomey, D., Jupyter for Data Science. 2017: Packt Publishing. ISBN-10: 1785880071 Váradi, C., C. Lew, and A. Guttman R,apid magnetic bead based sample preparation for automated and high throughput N-glycan analysis of therapeutic antibodies. Anal Chem, 2014. 86(12): p.5682-7. Wang, Y., et al., Comprehensive serum N-glycan profiling identifies a biomarker panel for early diagnosis of non-small-cell lung cancer, Proteomics 23(20) (2023), e2300140. Willers, H., et al., Basic mechanisms of therapeutic resistance to radiation and chemotherapy in lung cancer. Cancer J, 2013. 19(3): p.200-7. Yang, K.A.-O., et al., Economic burden of advanced lung cancer patients treated by gefitinib alone and combined with chemotherapy in two regions of China. J Med Econ. 2023 Jan- Dec;26(1):1424-1431.

Claims

Claims 1. A method for predicting the effectiveness ofm choetherapy treatment in a patient with cancer disease, preferably lung cancer disease, said mde ctohmoprising a.

1. providing a first serum sample of the pat tieanketn at time one, a.

2. providing a first serum sample of the pat tieanketn at time two, wherein said patient is treated by chemotherap thye in interval between time one and time two, b.

1. obtaining a first N-glycome analysis resuoltm fr the first serum sample, to obtain a first glycan structure features pattern, b.

2. obtaining a second N-glycome analysis resroumlt f the second serum sample, to obtain a second glycan structure features pattern, c. comparing the first glycan structure featurete prant and the second glycan structure feature pna ttoter assess changes in the patterns relating to starulct huarnges in the N-glycome of the patient, byy cinagrr out classification task to predict 'regression', 'perosgsrion' and 'stationary' states for characteri tzhieng respective chemotherapy outcome in the patient.

2. The method of claim 1, wherein the N-glycomely asnisa results are obtained by separation of N- glycans (preferably asparagine-linked glycans) fr thoem samples and glycan structure features argen aesdsi to the separated N-glycan peaks; preferably thely Nca-gns are separated by capillary electrophoresis measurement on the samples, and the glycan steru fcetautrures / characteristics are the relative are paea okfs of the separated and relevant N-glycans; 3. The method of claim 1 or 2, wherein the structu feral tures are the relative area of peaks of theseparated and relevant N-glycans.

4. The method of claims of any of 1 to 3, wherein N th-eglycome analysis is carried out by capillary electorphoresis list of the relevant glycans foer c thlassification task are selected by peak area 5. The method of any of claims 1 to 4, wherein N th-gelycan profile is correlated with the outcome of the treatment of a cancer.

6. The method of any of claims 1 to 5, wherein g thlyecans for the classification task are selecteodm fr the group consisting of the following N-linked galync structures: G1: FA4BG4[3,3,3,3]S4, G2: A2G2[6]S2, G3: FA3G3[6]S3, G4: A2G2[3]S2, G5: A2BG2S2, M3, G6: FA2G2S2, G7: FA2BG2S2,G8: FA2[6] G1S1, G9: A3G3[3]S2, G10: A2G2[6]S1, G11: A2BG2S1, G12: FA2G2S1, G13: FA2BG2S1, M7, G14: A4G4[6]S2, G15: FA2, M6, G16: FA2B, G17: FA2[6]G1, M7, G18: FA2[3]G1, G19: FA2B[6]G1, M8, G20: FA2G2, G21: M9 7. The method of any of claims 1 to 4, wherein l tihste of the relevant glycans for the binaryclassification tasks comprises at least 2, 3, or 45 glycans selected from the group consisting th oeffollowing glycans: G6, G12, G13, G20, G21.

8. The method of claim 6 or 7 wherein the listh oef t relevant glycans for the binary classificationtasks comprises at least 2, 3, 4 or 5 glycanst seedle frcom the group consisting of the following galyncs: for regression: G1, G2, G6, G12, G13, G14, G15, G17, G19, G20, G pr2e1f,erably G13, G14, G15, G17, G19, G20, G21, and / or for progression, G3, G4, G6, G8, G12, G13, G16, G20, G21, prefer farbolmy G3, G4, G6, G8, G12, G1 G3,20, G21, and / or for stationary status: G2, G3, G6, G11, G12, G13, G16, G18, G19, G20, G or21 p,referably from G3, G6, G11, G12, G13, G16, G18, G19, G21.

9. The method of any of claims 1 to 8, wherein s calaidssifier is trained with using independent data sets generated by randomly splitting the originaatal d set for each classification tasks.

10. The method of any of claims 1 to 9, wherein N th-gelycome (pattern) is correlated with the outcomeof the treatment of a cancer, preferably a canecler ct sed from the following group of cancers:preferably selected from the group consisting ofn- nsmoall cell lung cancers, preferably adenocarcinoma, squamous cell carcinoma, and c laerlgl cearcinoma.

11. The method of any of claims 9 to 10, whereein tr thained classifier is tested on independent s deattsa generated by randomly splitting the original daetta fo sr each classification task.

12. The method of claim 11, wherein a trained compu atlgeorrithm is used for classification and is carried out by a computer, preferably programme cda trory out a method according to any of claims 1 to 11, in particular at least step c. of claim 1.

13. The method of any of claims 1 to 12, wherein salaidss cifier algorithm comprises using machine learning or deep learning algorithms; wherein prraebfely said relevant glycans may undergo additional feature selection.

14. The method of any of claims 1 to 13, wherein trheea a under curve (AUC) values of the receiver operating characteristic (ROC) curves providedh bey c tlassifier are at least 0,8, preferably at l 0e,a9s.t

Citation Information

Patent Citations

  • Method for the analysis of n-glycans attached to immunoglobulin g from human blood plasma and its use

    US20160103137A1

  • Methods and products related to the improved analysis of carbohydrates

    US8209132B2

  • Clinical diagnostics using glycans

    WO2023034383A2