Composition, kit, medical sample classification method, device and storage medium

By using a composition containing decoy oligonucleotides and a machine learning model, combined with TERT promoter mutation and NRN1 methylation information, the problem of insufficient sensitivity in urine cytology was addressed, and the accuracy of urine sample detection was improved.

WO2026020325A9PCT designated stage Publication Date: 2026-04-09BOE TECHNOLOGY GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Urine cytology tests have low sensitivity and are prone to false negatives, resulting in low accuracy of test results.

Method used

Using a composition containing several different decoy oligonucleotides, gene mutations and methylation were detected by dPCR, and medical samples were classified by combining machine learning models. TERT promoter mutations and NRN1 methylation information were used for auxiliary diagnosis.

Benefits of technology

It improves the accuracy of medical sample classification, assisting medical personnel in making more accurate diagnoses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024107078_09042026_PF_FP_ABST
    Figure CN2024107078_09042026_PF_FP_ABST
Patent Text Reader

Abstract

A composition, a kit, a medical sample classification method, a device and a storage medium, which belong to the field of medical detection. The method comprises: acquiring input information of a first medical sample, wherein the input information of the first medical sample comprises gene measurement values of the first medical sample, and the gene measurement values are used for indicating at least one of a gene mutation status and a gene methylation status; inputting the input information of the first medical sample into a classification model to obtain an output result of the classification model, wherein the output result is used for indicating the probability distribution of the first medical sample across candidate classes, and the classification model is a machine learning model obtained by means of training based on input information of second medical samples and classes of the second medical samples; and on the basis of the output result, acquiring a classification result of the first medical sample, wherein the classification result is used for indicating the class of the first medical sample. The solution can improve the accuracy of classification results of first medical samples, and the classification results can assist medical staff in diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Composition, kit and method, apparatus, storage medium for classifying medical samples TECHNICAL FIELD

[0001] The present application relates to the technical field of medical detection, and particularly relates to a composition, a kit and a method, apparatus and storage medium for classifying medical samples. BACKGROUND

[0002] Urine sample detection is a common detection method for assisting medical personnel in diagnosing urinary system problems.

[0003] In related technologies, urine cytology examination is one of the means for diagnosing urinary system problems. Specifically, through urine cytology examination, information such as the content of abnormal cells in the urine sample can be determined, so that medical personnel can make a diagnosis by referring to the information such as the content of abnormal cells in the urine sample.

[0004] However, the sensitivity of the urine cytology examination in the above scheme is not high, and false negative results are prone to occur, resulting in low accuracy of the detection result of the urine sample.

[0005] SUMMARY

[0006] Embodiments of the present application provide a composition, a kit and a method, apparatus and storage medium for classifying medical samples; the technical solutions are as follows:

[0007] In one aspect, the present application provides a composition, which comprises: a plurality of different decoy oligonucleotides, wherein the plurality of different decoy oligonucleotides are configured to collectively hybridize to a plurality of DNA molecules derived from a target genomic region.

[0008] Wherein each genomic region in the target genomic region is differentially methylated and / or differentially mutated in at least one first sample compared to a second sample.

[0009] In another aspect, the present application provides a kit, which comprises the above-mentioned composition.

[0010] In one aspect, the present application provides a method for classifying medical samples, which comprises:

[0011] Obtaining input information of a first medical sample, wherein the input information of the first medical sample comprises a gene measurement value of the first medical sample; the gene measurement value is used to indicate at least one of a gene mutation condition and a gene methylation condition;

[0012] inputting the input information of the first medical sample into a classification model to obtain an output result of the classification model; the output result is used to indicate a probability distribution of the first medical sample in each candidate classification; the classification model is a machine learning model trained by input information of a second medical sample and a classification of the second medical sample;

[0013] obtaining a classification result of the first medical sample according to the output result; the classification result is used to indicate the classification of the first medical sample.

[0014] In another aspect, an embodiment of the present application provides a classification device of a medical sample, the device comprising:

[0015] an information obtaining module configured to obtain input information of a first medical sample, the input information of the first medical sample comprising a gene measurement value of the first medical sample; the gene measurement value is used to indicate at least one of a gene mutation condition and a gene methylation condition;

[0016] an inputting module configured to input the input information of the first medical sample into a classification model to obtain an output result of the classification model; the output result is used to indicate a probability distribution of the first medical sample in each candidate classification; the classification model is a machine learning model trained by input information of a second medical sample and a classification of the second medical sample;

[0017] a result obtaining module configured to obtain a classification result of the first medical sample according to the output result; the classification result is used to indicate the classification of the first medical sample.

[0018] In some embodiments, the gene measurement value comprises at least one of a first gene measurement value and a second gene measurement value; the first gene measurement value is used to indicate a mutation allele frequency MAF of a TERT promoter mutation, and the second gene measurement value is used to indicate a NRN1 methylation content.

[0019] In some embodiments, the input information of the first medical sample further comprises one or more of the following information:

[0020] attribute information of a user corresponding to the first medical sample;

[0021] a detection result of an abnormal cell in the first medical sample;

[0022] a medical image detection result of a tissue region corresponding to the first medical sample.

[0023] In some embodiments, the detection result of the abnormal cell in the first medical sample comprises one or more of the following information:

[0024] a proportion of abnormal cells in the first medical sample;

[0025] a classification of abnormal cells in the first medical sample.

[0026] In some embodiments, the medical image detection result of the tissue region corresponding to the first medical sample comprises one or more of the following information:

[0027] a classification of the tissue region corresponding to the first medical sample;

[0028] an area proportion of an abnormal region in the tissue region corresponding to the first medical sample;

[0029] position information of the abnormal region of the tissue region corresponding to the first medical sample;

[0030] a medical image of the tissue region corresponding to the first medical sample.

[0031] In some embodiments, the first medical sample is at least one of the following medical samples:

[0032] a urine sample, a blood sample, a tissue section sample.

[0033] In some embodiments, the input module is configured to multiply each of a plurality of data in the input information of the first medical sample by a respective weight of the plurality of data to obtain optimized input information.

[0034] The input module is configured to input the optimized input information into the classification model to obtain an output result of the classification model.

[0035] In some embodiments, the candidate classifications include a first candidate classification and a second candidate classification, the first candidate classification is a non-designated classification, and the second candidate classification is a designated classification; the classification of the second medical sample is one of the first candidate classification and the second candidate classification.

[0036] Alternatively,

[0037] The candidate classifications include a first candidate classification and at least two third candidate classifications, the first candidate classification is a non-designated classification, and the at least two third candidate classifications are at least two sub-classifications under a designated classification; the classification of the second medical sample is one of the first candidate classification and the at least two third candidate classifications.

[0038] In some embodiments, the apparatus further comprises a training module configured to input the input information of the second medical sample into the classification model to obtain a prediction result output by the classification model, wherein the prediction result is used to indicate a probability distribution of the second medical sample in the candidate classifications.

[0039] The training module is configured to update model parameters of the classification model according to a difference between the probability distribution of the second medical sample in the candidate classifications and a classification of the second medical sample.

[0040] In some embodiments, the classification model comprises a plurality of network layers, each of which has a respective weight matrix.

[0041] The model parameters of the classification model comprise:

[0042] The weight matrix of each of the plurality of network layers.

[0043] Or,

[0044] The weight matrix of each of the plurality of network layers, and a weight of each of a plurality of data in the input information.

[0045] In some embodiments, the gene measurement value of the first medical sample is obtained by performing dPCR detection on the first medical sample using a kit.

[0046] In some embodiments, the kit comprises a composition comprising a plurality of different decoy oligonucleotides, wherein the plurality of different decoy oligonucleotides are configured to collectively hybridize to a plurality of DNA molecules derived from a target genomic region.

[0047] Each of the target genomic regions is differentially methylated and / or differentially mutated in at least one first sample compared to a second sample.

[0048] In some embodiments, the mutation refers to a TERT promoter region mutation, and the TERT promoter region mutation comprises at least one of C228T and C250T.

[0049] In some embodiments, the plurality of DNA molecules are derived from at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic region of any one of sequences 1-2.

[0050] In some embodiments, the plurality of DNA molecules are derived from at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic region of SEQ ID NO: 1-2.

[0051] In another aspect, the embodiments of the present application provide a computer device, comprising a processor and a memory, wherein the memory stores at least one computer instruction, and the at least one computer instruction is loaded and executed by the processor to implement the classification method of the medical sample according to the above aspect.

[0052] In another aspect, the embodiments of the present application provide a computer readable storage medium, wherein the readable storage medium stores at least one computer instruction, and the at least one computer instruction is loaded and executed by a processor to implement the classification method of the medical sample according to the above aspect.

[0053] In yet another aspect, the embodiments of the present application further provide a computer program product, comprising computer instructions stored in a computer readable storage medium, wherein a processor reads and executes the computer instructions from the computer readable storage medium to implement the classification method of the medical sample according to the above aspect.

[0054] The embodiments of the present application provide a scheme, which takes the input information (such as gene measurement value) of the first medical sample as the classification basis, wherein the gene measurement value can indicate at least one of the gene mutation condition and the gene methylation condition, and the gene mutation condition and the gene methylation condition are closely related to the health condition of the human body; specifically, a classification model for auxiliary diagnosis is trained according to the input information of the second medical sample and the classification of the second medical sample, and in the application process, the input information of the first medical sample is input into the classification model to obtain the classification result of the first medical sample; the scheme introduces the machine learning model (i.e., the above classification model), which can predict the classification of the first medical sample according to at least one of the gene mutation condition and the gene methylation condition of the first medical sample, and thus the accuracy of the classification of the medical sample can be improved, thereby assisting medical personnel to make more accurate diagnosis. BRIEF DESCRIPTION OF DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0056] Fig. 1 is an architecture diagram of a computer system according to an example embodiment of the present application;

[0057] Fig. 2 is a flowchart of a medical sample classification method according to an example embodiment of the present application;

[0058] Fig. 3 is an implementation flowchart of a medical sample classification method according to an example embodiment of the present application;

[0059] Fig. 4 is a machine learning algorithm performance comparison diagram according to an example embodiment of the present application;

[0060] Fig. 5 is a detection weight optimization comparison diagram according to an example embodiment of the present application;

[0061] Fig. 6 is a clinical sample experiment flowchart according to an example embodiment of the present application;

[0062] Fig. 7 is a dynamic linear range detection result diagram of TERT promoter mutation according to an example embodiment of the present application;

[0063] Fig. 8 is a linear range test dPCR result diagram of TERT promoter mutation according to an example embodiment of the present application;

[0064] Fig. 9 is a dPCR linear range test result diagram of NRN1 methylation biomarker according to an example embodiment of the present application;

[0065] Fig. 10 is a dPCR test result distribution diagram according to an example embodiment of the present application;

[0066] Fig. 11 is a detection result distribution diagram according to an example embodiment of the present application;

[0067] Fig. 12 is a machine learning algorithm performance comparison diagram according to an example embodiment of the present application;

[0068] Fig. 13 is a ROC curve diagram and test set confusion matrix diagram of a prediction test set according to an example embodiment of the present application;

[0069] Fig. 14 is a block diagram of a medical sample classification apparatus according to an example embodiment of the present application;

[0070] Fig. 15 is a structural block diagram of a computer device according to an example embodiment of the present application. DETAILED DESCRIPTION

[0071] To make the purpose, technical scheme and advantages of the present application clearer, the embodiments of the present application will be described in further detail below with reference to the drawings.

[0072] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The description of the exemplary embodiments is intended to apply to any exemplary embodiment, unless specified otherwise. It is believed that the application will be better understood from the following description with reference to the drawings, in which:

[0073] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0074] It should be understood that although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a particular order or hierarchy. For example, a first parameter can be termed a second parameter, and, similarly, a second parameter can be termed a first parameter, without departing from the scope of the present disclosure. As used herein, the term "if' can be construed to mean "when" or "in response to determining" or "in response to a determination" that a certain condition precedent has been satisfied or obtained, unless and except the context clearly indicates otherwise.

[0075] Before introducing the technical solutions of the present application, some background technical knowledge related to the present application will be introduced and explained. The following related technologies can be combined with the technical solutions of the embodiments of the present application in any way as optional solutions, which all belong to the protection scope of the embodiments of the present application. The embodiments of the present application include at least part of the following contents:

[0076] 1) Digital PCR (dPCR): A highly sensitive and accurate molecular biology technique for absolute quantification of the copy number of a specific DNA sequence. Compared with quantitative PCR (qPCR), dPCR does not require a standard curve, because it directly divides the sample into thousands of micro-reaction units, each containing zero, one or more target molecules, and then determines the presence or absence of target DNA sequences in each unit through fluorescence signals or other detection methods, thereby directly calculating the absolute number of target sequences.

[0077] 2) Machine Learning: Machine Learning is a multidisciplinary field that focuses on how computers can learn from data and improve their performance on specific tasks without explicit programming. Machine Learning is a core part of Artificial Intelligence (AI), enabling systems to automatically identify patterns, mine valuable information from data, and make predictions or decisions based on that information. Supervised learning is the most common type of machine learning, where models learn from known input-output pairs (training data).

[0078] 3) Classification Model: Classification Model is an important branch of machine learning that aims to predict the class of the output based on the input data. This type of model has wide applications in many fields, such as medical diagnosis, financial risk assessment, image recognition, and text classification. Here are some common classification models and their characteristics:

[0079] Logistic Regression: Logistic Regression is a binary classification model that converts the continuous output of linear regression into a probability value between 0 and 1 using a non-linear (sigmoid) function, which then determines the likelihood of belonging to a certain class.

[0080] Support Vector Machines (SVM): Support Vector Machines find the optimal hyperplane to classify data, especially good at handling high-dimensional data and small sample situations, and can handle non-linear problems through kernel tricks.

[0081] Random Forests: Random Forests is an ensemble learning method based on decision trees, which improves the accuracy and stability of classification by constructing multiple decision trees and combining their prediction results.

[0082] The scheme provided by the embodiments of the present application relates to dPCR detection, classification model and other technologies, which are specifically explained by the following embodiments.

[0083] Please refer to FIG. 1, which shows the architecture diagram of a computer system provided by an exemplary embodiment of the present application. The computer system can realize the architecture of the configuration system of the medical sample classification method. As shown in FIG. 1, the computer system can include a terminal device 110, a server 120, and a test device 130. The terminal device 110, the server 120, and the test device 130 can be directly or indirectly connected by wired or wireless communication methods, which are not limited in the present application.

[0084] Optionally, the terminal device 110 can be a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal device 110 can install and run a client of a target application, which can be an application with natural language processing function, such as a generative language model (e.g., a large language model). The target application is not limited in form in the present application, including but not limited to an application (App) installed in the terminal device 110, a mini-program, etc., and can also be in the form of a web page.

[0085] Optionally, the server 120 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The cloud server of the big data and artificial intelligence platform can provide artificial intelligence cloud services. The server 120 can be a background server of the target application, configured to provide background services for the client of the target application.

[0086] Optionally, the test device 130 can be a polymerase chain reaction (PCR) device, a next generation sequencing (NGS) instrument, a gene chip scanner, a flow cytometer, a microwell plate instrument, an electrophoresis instrument, a nucleic acid extractor, etc., but is not limited thereto. The test device 130 is used in genomics, molecular biology, and genetics research, for gene sequence analysis, gene expression determination, genetic variation detection, etc. The test device 130 can provide detection services for the terminal device 110 and the server 120.

[0087] For example, the classification method of medical samples is applied to the scenario of medical personnel diagnosing urinary system problems.

[0088] The test device 130 performs dPCR detection on the first medical sample 101 to obtain a gene determination value 102 of the first medical sample 101, and sends the gene determination value 102 to the terminal device 110. The gene determination value 102 is used to indicate the gene mutation condition and the gene methylation condition.

[0089] In a case where the genetic measurement value 102 of the first medical sample 101 is acquired, the terminal device 110 can input the genetic measurement value 102 of the first medical sample 101 into the classification model 103 to obtain an output result of the classification model 103, the output result being used to indicate a probability distribution of the first medical sample in each candidate classification, the classification model 103 being a machine learning model trained by the genetic measurement value of the second medical sample and the classification of the second medical sample; according to the output result, the terminal device 110 can acquire the classification result 104 of the first medical sample 101; the classification result 104 is used to indicate the classification of the first medical sample 101. The classification model 103 can be provided by the server 120 to the terminal device 110.

[0090] In some embodiments, after the terminal device 110 acquires the genetic measurement value 102 of the first medical sample 101, the terminal device 110 can send the genetic measurement value 102 of the first medical sample 101 to the server 120; then, the server 120 can input the input information of the first medical sample into the classification model 103 to obtain an output result of the classification model 103; then, the server 120 can send the output result of the classification model 103 to the terminal device 110; finally, according to the output result, the terminal device 110 can acquire the classification result 104 of the first medical sample.

[0091] Please refer to FIG. 2, which shows a flowchart of a medical sample classification method according to an example embodiment of the present application. The method is executed by a computer device, which can be the terminal device 110 in the system shown in FIG. 1; or, it can also be the server 120, or it can also be the terminal device 110 and the server 120. As shown in FIG. 2, the method can include steps 210, 220 and 230.

[0092] Step 210: acquiring input information of a first medical sample, the input information of the first medical sample including a genetic measurement value of the first medical sample; the genetic measurement value being used to indicate at least one of a genetic mutation condition and a genetic methylation condition.

[0093] The first medical sample can be a biological sample (such as urine) collected from a research subject (such as a user to be confirmed whether there is a urinary system problem).

[0094] The genetic measurement value can be a numerical value used to indicate a genetic mutation condition, such as a genetic mutation content; or, the genetic measurement value can be a numerical value used to indicate a genetic methylation condition, such as a genetic methylation content; or, the genetic measurement value can be a numerical value used to indicate a genetic mutation condition and a genetic methylation condition.

[0095] Exemplarily, the gene measurement value can be obtained by at least one of the following gene measurement methods: double-deoxy chain termination method (Sanger) sequencing, next generation sequencing (NGS), polymerase chain reaction (PCR) and PCR derived technology (such as real-time fluorescent quantitative PCR, digital PCR, etc.), gene chip technology.

[0096] Step 220: inputting the input information of the first medical sample into the classification model to obtain an output result of the classification model; the output result is used to indicate a probability distribution of the first medical sample in each candidate classification; the classification model is a machine learning model trained by input information of a second medical sample and a classification of the second medical sample.

[0097] In the embodiment of the present application, the second medical sample and the first medical sample are medical samples of the same type.

[0098] Exemplarily, the output result of the classification model is a series of probability values, each probability value corresponding to each candidate classification. For example, if the classification model is used for classification of whether there is a urinary system problem, the output result of the classification model may show the probability of the sample belonging to normal, abnormal and other categories.

[0099] In the embodiment of the present application, the classification model is a machine learning model that has been trained based on a supervised learning algorithm. Exemplarily, the supervised learning algorithm can be at least one of the following algorithms: logistic regression, support vector machine, random forest, neural network, etc.

[0100] Step 230: obtaining a classification result of the first medical sample according to the output result; the classification result is used to indicate the classification of the first medical sample.

[0101] In the embodiment of the present application, the computer device can determine the candidate classification with the highest probability as the classification result of the first medical sample according to the probability distribution of the first medical sample in each candidate classification output by the classification model. For example, if the output of the classification model shows that the probability of the first medical sample in classification 1 is 0.15, the probability in classification 2 is 0.20, and the probability in classification 3 is 0.65, the computer device can determine that the classification result of the first medical sample is classification 3.

[0102] Exemplarily, in order to improve the confidence of the prediction of the classification model, the developer can set a probability threshold. When the probability of a certain classification is higher than the probability threshold, it is regarded as the final classification result. When the probability of a certain classification is lower than the probability threshold, it can be re-detected or further analyzed. For example, the probability threshold can be 80%, or 60%.

[0103] The classification result of the first medical sample can assist medical personnel in diagnosis. According to the classification result, the medical personnel can perform further medical intervention, laboratory detection, and the like.

[0104] To sum up, the scheme provided by the embodiments of the present application takes the input information (such as gene measurement value) of the first medical sample as the classification basis, wherein the gene measurement value can indicate at least one of the gene mutation condition and the gene methylation condition, and the gene mutation condition and the gene methylation condition are closely related to the health condition of the human body. Specifically, a classification model for assisting diagnosis is trained according to the input information of the second medical sample and the second medical sample, and in the application process, the input information of the first medical sample is input into the classification model to obtain the classification result of the first medical sample. The scheme introduces a machine learning model (i.e., the above-mentioned classification model), which can predict the classification of the first medical sample according to at least one of the gene mutation condition and the gene methylation condition of the first medical sample, thereby improving the accuracy of the classification of the medical sample, and assisting medical personnel to make more accurate diagnosis.

[0105] Based on the scheme in the embodiment shown in FIG. 2, in a possible implementation, the gene measurement value includes at least one of a first gene measurement value and a second gene measurement value; the first gene measurement value is used to indicate the mutation allele frequency MAF of the TERT promoter mutation, and the second gene measurement value is used to indicate the methylation content of NRN1.

[0106] The first gene measurement value corresponds to the mutation analysis result of the telomerase reverse transcriptase (TERT) promoter region, in particular, the mutation allele frequency (MAF). The TERT promoter mutation is common in urinary system problems and the like, and the TERT promoter mutation can affect the expression of the TERT gene. The mutation allele frequency (MAF) is an index for measuring the proportion of a specific mutation in a cell population, which can be expressed in percentage. For example, if the MAF of the TERT promoter mutation is 20%, it means that 20% of the DNA molecules in the sequenced sample carry mutations at that position.

[0107] The second gene measurement value corresponds to the methylation content of the Neuregulin 1 (NRN1) gene. Gene methylation is an important epigenetic modification that can affect gene expression without changing the DNA sequence itself. The change in the methylation level of NRN1, as a specific gene, can be related to urinary system problems and the like. By measuring the methylation content of the NRN1 gene, important information for understanding urinary system problems and the like can be provided.

[0108] Based on the above embodiments, the present embodiment shows the specific content of the gene measurement value, the first gene measurement value and the second gene measurement value can be used in combination or individually, and the first gene measurement value and the second gene measurement value can reveal the biological state of the first medical sample from different angles. The TERT promoter mutation reflects the direct change of the gene sequence, while the methylation of NRN1 provides information at the gene expression regulation level. Integrating the first gene measurement value and the second gene measurement value into the analysis model can improve the accuracy of the classification result of the first medical sample, thereby better assisting the diagnosis of medical personnel.

[0109] Taking the application of the above medical sample classification method to the detection of urinary system problems as an example, please refer to FIG. 3, which shows the implementation flowchart of the medical sample classification method provided by an exemplary embodiment of the present application. As shown in FIG. 3, the present method can distinguish between medical samples with urinary system problems and normal medical samples in the first medical sample. First, the biomarker of the first medical sample (i.e. the input information described above) is obtained, which includes the TERT mutation content and the NRN1 methylation content; second, the TERT mutation content (i.e. X1 in FIG. 3) and the NRN1 methylation content (i.e. X2 in FIG. 3) of the first medical sample are input into the classification model; finally, the classification model can output the classification result of the first medical sample.

[0110] Among them, the above classification model can be obtained by training the TERT mutation content and the NRN1 methylation content of the second medical sample, and the classification of the second medical sample (the training steps of the classification model are described above). Exemplarily, data preprocessing can be performed on the TERT mutation content and the NRN1 methylation content of the second medical sample. The data preprocessing can be a visualization processing of the TERT mutation content and the NRN1 methylation content of the second medical sample, and the second medical sample is divided into a training set and a test set; then, the TERT mutation content and the NRN1 methylation content of the second medical sample in the training set are input into the classification model to obtain the prediction result output by the classification model, and the model parameters of the classification model are updated according to the difference between the probability distribution of the second medical sample in each candidate classification and the classification of the second medical sample; finally, the accuracy of the classification model is verified by the test set.

[0111] In addition, the above model parameters also include the weight value of the TERT mutation content (i.e. W1 in FIG. 3) and the weight value of the NRN1 methylation content (i.e. W2 in FIG. 3).

[0112] For example, the second medical samples include medical samples of patients with urinary system problems and medical samples of normal persons. For example, the second medical samples have a total of 77, the medical samples of patients with urinary system problems have 44, and the medical samples of normal persons have 35, of which 24 are stage I patients and 17 are stage II+ patients. When distinguishing between normal persons and patients with urinary system problems, all samples are randomly divided into a training set and a test set according to a ratio of 6:4, wherein the training set has 46 samples and the test set has 31 samples. When distinguishing between patient stages (normal person NC, stage I patient, and stage II+ patient), all samples are randomly divided into a training set and a test set according to a ratio of 5:5. The random number seed is 100.

[0113] For the three features of TERT, NRN1, and TERT+NRN1, three machine learning classification algorithms, namely, logistic regression (LR), support vector machine (SVM), and random forest (RF), are used respectively. The hyperparameters of each algorithm model are determined by combining 3-fold cross-validation and grid search algorithm on the training set; then, the AUC values of the three machine learning classification algorithms on the test set are compared. The performance comparison of the three machine learning classification algorithms is shown in FIG. 4, and it can be found that the performance of the classification model of the TERT+NRN1 combined feature is better than that of the classification model of a single TERT or NRN1 target feature; on the binary classification task, the AUC value of the logistic regression LR model is the highest.

[0114] Finally, in order to further measure the classification performance of the best model (i.e., the logistic regression LR model), the detection weights of each biomarker can be optimized, and the AUC graph of the specific detection result is shown in FIG. 5, from which it can be seen that the performance of the diagnosis model after weight ratio optimization is better than that of the original additive result diagnosis model. From the test set results, the accuracy of the optimized diagnosis model for distinguishing between normal and patient groups is as high as 93%.

[0115] In the case where the number of first medical samples is multiple, the TERT mutation content and the NRN1 methylation content of the multiple first medical samples can be input into the above classification model in batches and separately to obtain the classification results of each first medical sample; or the TERT mutation content and the NRN1 methylation content of the multiple first medical samples can be input into the above classification model at the same time to obtain the classification results of the multiple first medical samples.

[0116] Based on the schemes shown in the above various embodiments of the present application, in one possible implementation scheme, the input information of the first medical sample further includes one or more of the following information:

[0117] Attribute information of a user corresponding to the first medical sample;

[0118] Detection result of an abnormal cell in the first medical sample;

[0119] a medical image detection result of a tissue region corresponding to the first medical sample.

[0120] The user attribute information can include, but is not limited to, one or more of the following: basic information of the user, medical history, living habits, clinical symptoms, physical examination indicators, genetic information.

[0121] Specifically, the basic information of the user can include age, gender, etc.; the medical history of the user can include previous medical history, surgical history, family medical history, drug allergy history, vaccination record, etc.; the living habits of the user can include smoking history, drinking habits, eating habits, exercise frequency, etc.; the clinical symptoms of the user can include the main symptoms of the current visit, the time of the symptoms, the severity of the symptoms, etc.; the physical examination indicators of the user can include blood pressure, heart rate, weight, height, body mass index (BMI), blood routine, biochemical indicators, etc.; the genetic information of the user can include known genetic predisposition or genetic disease related information.

[0122] The abnormal cell detection result can be obtained by cytological detection technology. Specifically, the cytological detection technology can be flow cytometry, microscopic observation, immunohistochemical staining, etc.

[0123] The type of abnormal cell can be a specific abnormal cell type, such as a cancer cell, an inflammatory cell, an abnormal lymphocyte, etc.

[0124] The medical image detection result can be obtained by imaging examination technology. Specifically, the imaging examination technology can be computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography-computed tomography (PET-CT), ultrasound, etc.

[0125] The medical image detection result can indicate the abnormal structure or change of the abnormal region. Specifically, the medical image detection result can indicate the size, shape, boundary definition, enhancement characteristics, etc. of the abnormal region (such as a mass).

[0126] Based on the above embodiments, the present embodiment provides multiple optional contents of the input information of the first medical sample, and takes the attribute information of the user corresponding to the first medical sample, the detection result of the abnormal cells in the first medical sample, and the medical image detection result of the tissue region corresponding to the first medical sample as the input information, so as to further consider the influence of the attribute information of the user corresponding to the first medical sample, the detection result of the abnormal cells in the first medical sample, and the medical image detection result of the tissue region corresponding to the first medical sample on the classification result of the first medical sample, and further improve the accuracy of the classification result of the first medical sample.

[0127] Based on the schemes shown in the above embodiments of the present application, in a possible implementation, the detection result of the abnormal cells in the first medical sample includes one or more of the following information:

[0128] The proportion of the abnormal cells in the first medical sample;

[0129] The classification of the abnormal cells in the first medical sample.

[0130] The proportion of the abnormal cells refers to the proportion of the abnormal cells in the detected cell population. The proportion of the abnormal cells can be expressed by percentage, for example, if the detection result shows that the proportion of the abnormal cells is 15%, it means that about 15% of the cells in the first medical sample show abnormal characteristics, which is different from normal cells; this value can assist the judgment of medical personnel.

[0131] The classification of the abnormal cells can be classified according to the morphology, functional abnormality or specific pathological changes of the abnormal cells. For example, the classification of the abnormal cells can include but is not limited to one or more of the following: cancer cells, atypical cells, inflammatory cells, abnormal lymphocytes, macrophages / phagocytes.

[0132] Based on the above embodiments, the present embodiment proposes the specific content of the detection result of the abnormal cells in the first medical sample, which can specifically include the proportion of the abnormal cells and / or the classification of the abnormal cells, so as to further consider the influence of the proportion of the abnormal cells and / or the classification of the abnormal cells on the classification result of the first medical sample, and further improve the accuracy of the classification result of the first medical sample, and better assist the judgment of medical personnel.

[0133] Based on the schemes shown in the above embodiments of the present application, in a possible implementation, the medical image detection result of the tissue region corresponding to the first medical sample includes one or more of the following information:

[0134] The classification of the tissue region corresponding to the first medical sample;

[0135] an area proportion of the abnormal region in the tissue region corresponding to the first medical sample;

[0136] position information of the abnormal region of the tissue region corresponding to the first medical sample;

[0137] a medical image of the tissue region corresponding to the first medical sample.

[0138] The classification of the tissue region refers to the specific classification of the tissue or organ identified by medical imaging technology (such as ultrasound). Specifically, it can be indicated that the tissue region is part of the liver, lungs, brain, or other specific organs.

[0139] The area proportion of the abnormal region refers to the proportion of the identified abnormal region in the entire observed tissue region or organ. Specifically, the area proportion of the abnormal region can be expressed in percentage.

[0140] The position information of the abnormal region refers to the precise location of the abnormal region within the tissue or organ. Specifically, the position information of the abnormal region can include but is not limited to anatomical levels, distances from reference points, adjacent structural relationships, etc.

[0141] The medical image is a core part of directly displaying the condition of the tissue region, including image pictures, video materials, etc. Medical images not only show the anatomical structure of the tissue, but also reveal the functional state and pathological changes of the tissue through different imaging techniques.

[0142] Based on the above embodiments, this embodiment specifically introduces the specific content of the medical image detection result of the tissue region corresponding to the first medical sample, which can further consider the influence of the classification of the tissue region in the medical image detection result, the area proportion of the abnormal region, the position information of the abnormal region, and the medical image of the tissue region on the classification result of the first medical sample, further improving the accuracy of the classification result of the first medical sample, and better assisting medical personnel in judgment.

[0143] Based on the schemes shown in the above embodiments of the present application, in one possible implementation, the first medical sample is at least one of the following medical samples:

[0144] urine sample, blood sample, tissue section sample.

[0145] Correspondingly, the second medical sample is also at least one of the following medical samples: urine sample, blood sample, tissue section sample.

[0146] The urine sample is obtained by a non-invasive manner, which can reduce the damage to the user's body. The blood sample is a common sample for diagnosing various diseases and health conditions, and can be obtained by blood drawing. The tissue slice sample can be obtained by biopsy or surgery.

[0147] Based on the above embodiments, the first medical sample involves at least one of a urine sample, a blood sample, and a tissue slice sample, so that the present scheme can be applied to different types of samples such as urine samples, blood samples, and tissue slice samples, thereby expanding the application scenarios of the present scheme.

[0148] Based on the schemes shown in the above embodiments of the present application, in one possible implementation, the step 220 can be implemented as:

[0149] The plurality of data in the input information of the first medical sample is multiplied by the respective weight of the plurality of data to obtain optimized input information;

[0150] The optimized input information is input into the classification model to obtain an output result of the classification model.

[0151] In the embodiments of the present application, the greater the weight of the data, the more critical or more predictive the data is to the classification task. For example, the proportion of abnormal cells can be given a higher weight because it is directly related to urinary system problems.

[0152] The respective weight of the plurality of data can be set by the developer or can be the training result of the classification model. For example, the developer can set the respective weight of the plurality of data in the input information according to the correlation between the plurality of data and the output result. For example, the model parameters of the classification model can include the respective weight of the plurality of data.

[0153] For example, if the weight of the proportion of abnormal cells is 0.8 and the actual proportion is 15%, the optimized input information is 15% x 0.8 = 12%.

[0154] Based on the above embodiments, the present embodiment proposes a scheme of optimizing input information by using a weighting strategy. The data with a higher correlation with the output result has a higher weight, and thus the present scheme can improve the accuracy and robustness of the prediction of the classification model.

[0155] Based on the schemes shown in the above embodiments of the present application, in one possible implementation, each candidate classification includes a first candidate classification and a second candidate classification, the first candidate classification is a non-designated classification, and the second candidate classification is a designated classification; the classification of the second medical sample is one of the first candidate classification and the second candidate classification.

[0156] Alternatively,

[0157] The candidate classifications include a first candidate classification and at least two third candidate classifications, the first candidate classification being a non-designated classification, and the at least two third candidate classifications being at least two sub-classifications under a designated classification; and the classification of the second medical sample is one of the first candidate classification and the at least two third candidate classifications.

[0158] The first candidate classification is a non-designated classification, which means that the first candidate classification is a broad and non-specific classification, and the first candidate classification can be a default classification or another classification. The second candidate classification is a designated classification, which means that the second candidate classification is a specific and well-defined classification, and is a target for detailed classification. The classification of the second medical sample is one of the first candidate classification and the second candidate classification, which means that the second medical sample is classified into the broad first candidate classification or is accurately assigned to the designated second candidate classification.

[0159] For example, in the case of detecting urinary system problems, the first candidate classification can correspond to the classification result of the first medical sample without urinary system problems, and the second candidate classification can correspond to the classification result of the first medical sample with urinary system problems. The classification of the second medical sample can indicate whether the user of the second medical sample has urinary system problems.

[0160] The third candidate classification is at least two sub-classifications under a designated classification, which means that the third candidate classification is more detailed and specific.

[0161] For example, in the case of detecting urinary system problems, the first candidate classification can correspond to the classification result of the first medical sample without urinary system problems, and the third candidate classification can correspond to the classification result of the first medical sample with mild or severe urinary system problems.

[0162] In the embodiments of the present application, the classification model can be implemented as a binary classification model, or as a multi-classification model such as a three-classification model, a four-classification model, etc.

[0163] Based on the above embodiments, the present embodiment proposes a flexible classification process that can cover samples that are difficult to classify (through the first candidate classification), and can ensure fine classification of samples with clear characteristics (through the second or third candidate classification), thereby improving the flexibility of the application of the present scheme.

[0164] Based on the schemes shown in the above embodiments of the present application, in one possible implementation, the computer device can train the classification model before step 210. For example, the training steps of the classification model include:

[0165] Step 202: inputting the input information of the second medical sample into the classification model to obtain a prediction result output by the classification model; the prediction result is used to indicate a probability distribution of the second medical sample in each candidate classification;

[0166] Step 204: updating the model parameters of the classification model according to a difference between the probability distribution of the second medical sample in each candidate classification and the classification of the second medical sample.

[0167] For example, when the model parameters of the classification model are updated according to the difference between the probability distribution of the second medical sample in each candidate classification and the classification of the second medical sample, the computer device can calculate a loss function value according to the difference between the probability distribution in each candidate classification and the classification of the second medical sample, and update the model parameters of the classification model according to the loss function value, such as updating the model parameters of the classification model by means of reverse gradient propagation according to the loss function value.

[0168] The second medical sample can be a biological sample (such as urine) collected from a research subject. The research subject can include patients with urinary system problems and normal people without urinary system problems. The patients with urinary system problems can include patients with different levels of urinary system problems.

[0169] The input information of the second medical sample includes a gene measurement value of the second medical sample; the gene measurement value is used to indicate at least one of a gene mutation condition and a gene methylation condition. For example, the gene measurement value includes at least one of a first gene measurement value and a second gene measurement value.

[0170] After the computer device inputs the input information of the second medical sample into the classification model, the classification model can output the probability distribution of the second medical sample in each candidate classification according to the input information of the second medical sample. Since the classification of the second medical sample is known, the computer device can update the model parameters of the classification model according to the prediction result output by the classification model and the classification of the second medical sample, so as to reduce the prediction error of the classification model and further improve the accuracy of the prediction result output by the classification model.

[0171] Based on the schemes shown in the above various embodiments of the present application, in one possible implementation, the classification model includes a plurality of network layers, and each of the plurality of network layers has a respective weight matrix; the model parameters of the classification model include:

[0172] the respective weight matrix of each of the plurality of network layers;

[0173] or,

[0174] the respective weight matrix of each of the plurality of network layers, and the respective weight of each item of data in the input information.

[0175] wherein the classification model comprises a plurality of network layers, means that the classification model adopts a multi-layer neural network structure to more effectively process complex data sets and feature spaces. Multi-layer neural networks, or deep neural networks (DNNs), include but are not limited to convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTMs), and the like.

[0176] For example, the classification model can comprise an input layer, hidden layers, and an output layer. The input layer is used to receive the original data input to the classification model, i.e., the input information of the first medical sample, the input information of the second medical sample. The hidden layers are located between the input layer and the output layer and are used to learn the abstract representation of the data. Each hidden layer comprises a plurality of neurons, each of which has its own weight and bias parameters. The hidden layers can include convolutional layers, fully connected layers, normalization layers, pooling layers, recurrent layers, and the like.

[0177] The convolutional layer is used to detect local features in the original data. Each neuron of the fully connected layer is connected to all neurons of the previous layer and is used to integrate information from the previous layer, which is often used in the last few layers of the classification model. The normalization layer is used to accelerate the training process and improve the stability of the model. The pooling layer is used to reduce the spatial dimension of the data, reduce the amount of calculation, and retain important features. The recurrent layer is used to capture time or sequence dependencies when processing sequential data. The output layer is used to produce the final output result of the classification model, which can be a probability distribution indicating the likelihood of the second medical sample belonging to each candidate classification. For example, the activation function of the output layer can convert the output of the neuron to a probability value.

[0178] wherein the weight matrix of each network layer determines how the input information is converted into the output result by the network layer. For example, a convolutional neural network (CNN), a recurrent neural network (RNN), or a long short-term memory network (LSTM) can be composed of multiple hidden layers. Each layer has its own weight matrix, which determines how the input data is converted into the output and the way information is transmitted between layers. The size of the weight matrix depends on the input dimension and the output dimension of the layer. For example, for a fully connected layer, if the previous layer has m neurons and the current layer has n neurons, the size of the weight matrix will be mxn.

[0179] In addition, in addition to the weight matrix of each network layer, the classification model can also assign different weights to multiple data in the input information. For example, when multiple data in the input information contribute differently to the classification result, the model parameters of the classification model also include the weights of the multiple data in the input information.

[0180] During the training process, the initialization of the model parameters is usually random, and then the iterative optimization can be performed through the gradient descent algorithm and the backpropagation process to find the parameter values that minimize the loss function. This optimization process involves calculating the gradient of the loss function with respect to each weight and updating the weight by a certain learning rate.

[0181] Based on the above embodiments, the model parameters of the classification model can include the weight matrix of each network layer, or the weight matrix of each network layer and the weight of multiple data in the input information. During the training of the classification model, the model parameters are crucial to improving the performance and generalization ability of the classification model. The present scheme provides multiple options for model parameters, and the present scheme can reduce the prediction error of the classification model according to the influence degree of the weight matrix of the network layer and the weight of the multiple data on the output result, and further improve the accuracy of the prediction result output by the classification model.

[0182] Based on the schemes shown in the above embodiments of the present application, in one possible implementation, the gene measurement value of the first medical sample is a gene measurement value obtained by performing dPCR detection on the first medical sample by using a kit.

[0183] For example, the kit can be introduced into the dPCR chip array microwells by injection, and further introduced into each reaction microwell separated by dPCR oil (Bio-Rad). After sealing, PCR reaction is performed under the following conditions: 96℃ for 10 minutes, 96℃ for 50 cycles, each cycle for 30 seconds (slope 2.5 / second), 55℃ for 30 seconds, then 60℃ for 10 minutes, and finally kept at 10℃.

[0184] Among them, dPCR is an advanced molecular diagnostic technology, which has the following advantages compared with traditional qPCR:

[0185] Absolute quantification: dPCR can provide absolute molecular count, not relative quantification; dPCR technology divides the sample into thousands of independent reaction chambers, directly counts the number of target molecules, without relying on external standard curve, thereby improving the accuracy of the results;

[0186] High sensitivity: dPCR can detect extremely low levels of target sequences, even a single copy of DNA or RNA molecules, which makes it particularly useful in the detection of trace amounts of nucleic acids;

[0187] High precision and repeatability: Due to its digital nature, dPCR has a high tolerance for experimental variation, maintaining high consistency and precision even between different batches or operators, making it suitable for research and clinical applications that require extremely high accuracy;

[0188] Strong anti-interference ability: dPCR has good tolerance to inhibitors, allowing accurate detection of target sequences in complex samples containing PCR inhibitors such as blood and feces, reducing the burden of sample pretreatment;

[0189] Wide dynamic range: Although dPCR is usually used for low-abundance target detection, it actually has a wide dynamic range and can accurately detect target molecules over multiple orders of magnitude, from extremely low to relatively high copy numbers;

[0190] High specificity: By designing specific primers and probes, dPCR can effectively distinguish between highly similar sequences, reducing false positive results, which is important for identifying gene mutations, viral subtypes, etc.;

[0191] Wide applicability: dPCR is not only suitable for gene expression analysis and pathogen detection, but also widely used in fields such as gene mutation detection, copy number variation analysis, and methylation research.

[0192] For example, the above-mentioned dPCR detection can be implemented on a variety of different platforms, and the dPCR detection platform can be any one of an array microchamber type, a droplet type, and a microcavity type.

[0193] Among them, the array microchamber type (Microarray Chamber-based dPCR) platform uses microfabrication technology to manufacture thousands of tiny reaction chambers or microholes on a solid carrier. Each microhole corresponds to a separate PCR reaction container and can capture and amplify target nucleic acid molecules. The number of molecules in each microhole is determined by fluorescence or other detection methods, and the absolute copy number of the target sequence is directly calculated. The advantage of this method is high throughput and high automation.

[0194] Among them, the droplet type (Droplet Digital PCR, dPCR): The droplet dPCR platform uses microfluidic technology to divide samples and reaction mixtures into thousands of nanoliter-level microdroplets, each of which becomes an independent PCR reaction unit. After PCR amplification in these microdroplets, the presence or absence of target molecules is determined by detecting the fluorescence signal in each microdroplet. The technology of RainDance and other companies is based on this droplet microfluidic principle. Droplet dPCR is widely used in scientific research and clinical detection due to its high sensitivity and good reproducibility.

[0195] Microchamber or Microreactor-based dPCR: refers to any system that utilizes small, enclosed spaces (microchambers) for independent PCR reactions. These microchambers can be formed through various technological means, such as microfluidic chip technology, and they are also capable of achieving single-molecule level absolute quantification of nucleic acids. Similar to array-based micro-well systems, but may place more emphasis on the special design of microchamber structures to optimize reaction efficiency or detection sensitivity.

[0196] Based on the above embodiments, this embodiment proposes a feasible scheme for obtaining the gene measurement value of the first medical sample. Specifically, the gene measurement value of the first medical sample can be obtained by performing dPCR detection on the first medical sample by using a kit containing a specified composition. dPCR has the advantages of absolute quantification, high sensitivity, accuracy, and wide applicability, and can improve the accuracy of the gene measurement value of the first medical sample.

[0197] The embodiment of the present application provides a composition, which comprises: a plurality of different decoy oligonucleotides, wherein the plurality of different decoy oligonucleotides are configured to collectively hybridize to a plurality of DNA molecules derived from a target genomic region.

[0198] Wherein each genomic region in the target genomic region is differentially methylated and / or differentially mutated in at least one first sample compared to a second sample.

[0199] In the kit for performing dPCR detection on the first medical sample to obtain the gene measurement value of the first medical sample, the composition described above can be included.

[0200] In the embodiment of the present application, the first sample and the first medical sample involved in each of the above embodiments are the same type of medical sample, such as at least one of a urine sample, a blood sample, and a tissue section sample; the second sample and the first medical sample involved in each of the above embodiments are the same type of medical sample, such as at least one of a urine sample, a blood sample, and a tissue section sample.

[0201] Wherein the differential methylation means that the DNA methylation pattern of the target genomic region is different between the first sample and the second sample. Methylation is a form of epigenetic modification that can affect gene expression, so this difference may be related to changes in gene regulation.

[0202] Wherein the differential mutation means that there are significant variations in the DNA sequence of the target region in the two samples. These mutations can be point mutations, insertions, deletions, or more complex genomic structural changes, which can affect gene function or protein products.

[0203] In the experiment, the decoy oligonucleotide described above can be added to the mixture containing the target DNA molecules. Since the decoy oligonucleotide is complementary to the target DNA sequence, the decoy oligonucleotide can hybridize to the target DNA fragment by the principle of base pairing (such as adenine A pairing with thymine T, cytosine C pairing with guanine G). Through the subsequent washing step, the unbound DNA fragments are removed, and the DNA fragments bound to the decoy are retained, achieving the enrichment of the target region. The enriched DNA can be further sequenced to analyze the sequence information of the target region in depth, such as finding mutations, methylation status, or gene expression levels.

[0204] The present embodiment proposes a composition which can be used in high-throughput sequencing experiments, such as targeted sequencing, to study the methylation and mutation status of specific genomic regions, and can be used in the dPCR detection in the above embodiments to obtain the genetic determination value of the first medical sample.

[0205] Based on the scheme shown in the above various embodiments of the present application, in a possible implementation scheme, the mutation refers to a TERT promoter region mutation, and the TERT promoter region mutation includes at least one of C228T and C250T.

[0206] For example, taking urinary system problems as an example, two mutations are found in the TERT promoter, namely g.1295228C>T and g.1295250C>T, which are prevalent in 59-77% of urinary system problems. Therefore, the above composition can be used to detect at least one TERT promoter region mutation of C228T and C250T to assist medical personnel in diagnosis.

[0207] Based on the scheme shown in the above various embodiments of the present application, in a possible implementation scheme, the plurality of DNA molecules are derived from at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic regions of any one of sequences 1-2.

[0208] Based on the scheme shown in the above various embodiments of the present application, in a possible implementation scheme, the plurality of DNA molecules are derived from at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic regions of sequences 1-2.

[0209] Each of the target genomic regions of sequence 1 is differentially mutated in at least one first sample compared to a second sample. For example, the target genomic regions of sequence 1 can be:

[0210] wherein each of the target genomic regions of Sequence 2 is differentially methylated in at least one first sample compared to a second sample. Exemplarily, the target genomic regions of Sequence 2 can be:

[0211] wherein the bait oligonucleotide is a short synthetic DNA fragment with a sequence designed to be complementary to a specific target DNA sequence. Exemplarily, the bait oligonucleotide can be a primer set and a probe for identifying the target genomic region.

[0212] Please refer to Table 1, which shows a primer set and a probe provided by an exemplary embodiment of the present application, which can be applied to the detection of TERT promoter mutation C228T, and the primer and probe sequences relative to the wild type; the primer set includes a forward primer (fw_primer) and a reverse primer (rev_primer), the nucleic acid sequence of the forward primer is ACCCCTCCCGGGTC, and the nucleic acid sequence of the reverse primer is CCGCGGAAAGGAAG; the probe includes a wild type probe (wt_probe) and a mutant probe (mut_probe), the nucleic acid sequence of the wild type probe is AGGGCCCGGAgGGGGCTGGG, and the nucleic acid sequence of the mutant probe is AGGGCCCGGAaGGGGCTGGG.

[0213] Table 1

[0214] wherein FAM_IowaBlack and HEX_IowaBlack in Table 1 above are commonly used fluorescently labeled probe systems in quantitative PCR (qPCR).

[0215] FAM in FAM_IowaBlack represents Carboxyfluorescein, which is a green fluorescent dye commonly used to label probes in PCR reactions. When PCR amplification occurs, the FAM-labeled probe pairs with a specific DNA sequence, and as the DNA strand is extended, it is cut by an enzyme, releasing the FAM dye, which emits a fluorescent signal, which can be used to monitor the amplification of the target gene in real time. IowaBlack refers to a black quencher, which is tightly attached to the other end of the probe, effectively quenching the fluorescence on the uncut probe, thereby improving the specificity and sensitivity of the signal. Therefore, FAM_IowaBlack refers to a probe system with FAM fluorescent labeling and IowaBlack as a quencher.

[0216] HEX in HEX_IowaBlack stands for Hexachloro-6-carboxyfluorescein, which is a similar but slightly different wavelength fluorescent dye that typically emits yellow or orange fluorescence. Functionally similar to FAM, HEX is also used to label probes to track different target sequences in qPCR reactions. Similar to FAM_IowaBlack probes, HEX_IowaBlack probes combine the HEX fluorescent dye and IowaBlack quencher to ensure efficient, specific signal generation for differentiating between different target sequences in multiplex PCR assays.

[0217] That is, FAM_IowaBlack and HEX_IowaBlack are two different fluorescently labeled probe systems, each carrying a specific fluorescent dye (FAM or HEX) and IowaBlack quencher, widely used in fields such as gene expression analysis, pathogen detection, and genetic variation research.

[0218] Please refer to Table 2, which shows the primer set and probe provided by an exemplary embodiment of the present application, which can be applied to the detection of TERT promoter mutation C250T and the primer probe sequence relative to the wild type; the primer set includes a forward primer and a reverse primer, the nucleic acid sequence of the forward primer is CTTCCAGCTCCGCCTCCT, and the nucleic acid sequence of the reverse primer is GGCCGCGGAAAGGAAGGG; the probe includes a wild type probe and a mutant probe, the nucleic acid sequence of the wild type probe is TCCCGACCCCTcCCGGGTCC, and the nucleic acid sequence of the mutant probe is TCCCGACCCCTtCCGGGTCC.

[0219] Table 2

[0220] Please refer to Table 3, which shows the primer set and probe provided by an exemplary embodiment of the present application, which can be applied to the detection of NRN1 methylation hotspot; the primer set includes a forward primer and a reverse primer, the nucleic acid sequence of the forward primer is TTTTAAAATTTAGTGGYGAAAGAAGC, and the nucleic acid sequence of the reverse primer is TTTTACACAGCACACACTCACACATGC; the nucleic acid sequence of the probe is TTTAAAATGCATTGATATTTTTTAAGGGG.

[0221] Table 3

[0222] In summary, the composition shown in the embodiments of the present application can be applied to the primer set and probe for identifying the target genomic region, and the dPCR detection method developed based on the composition has the advantages of high sensitivity, strong specificity, rapid reaction, low requirement for instrument equipment, etc., and can improve the accuracy of dPCR detection, especially the accuracy of dPCR detection of medical samples of the urinary system.

[0223] Based on the schemes shown in the above various embodiments of the present application, in one possible implementation, a kit includes the above composition.

[0224] For example, the kit can include template DNA, 2x ddPCR supermix (Bio-Rad), forward and reverse primers (9 micromoles each), FAM and HEX fluorescent probes for mutation detection (2.5 micromoles for each mutation, respectively).

[0225] The kit shown in the embodiments of the present application can be applied to the reaction mixture of dPCR detection, and the dPCR detection method developed based on the composition has the advantages of high sensitivity, strong specificity, rapid reaction, low requirement for instrument equipment, etc., and can improve the accuracy of dPCR detection, especially the accuracy of dPCR detection of medical samples of the urinary system.

[0226] At present, the gold standard for detecting bladder cancer is the expensive and invasive cystoscopy method, which is not conducive to large-scale screening of high-risk populations. Due to the high recurrence characteristics of non-muscular invasive bladder cancer, bladder cancer patients need to undergo regular cystoscopy detection after surgery / chemotherapy intervention, which increases the economic burden and physical damage of patients and may cause other complications. Other clinical detection methods such as urine cytology have the problem of low sensitivity, making it difficult to apply to direct detection.

[0227] For example, based on the scheme shown in any one or more of the above embodiments of the present application, the embodiments of the present application propose a high-efficiency non-invasive bladder cancer detection method. The method can detect the mutation allele frequency (MAF) of TERT promoter region mutations (C228T and C250T) and NRN1 methylation characteristic sites as combined biomarkers by dPCR at the same time, and then use machine learning algorithm to optimize the weight proportion of each biomarker, thereby realizing accurate detection of bladder cancer and to a certain extent, realizing staging detection of bladder cancer.

[0228] Bladder Cancer (BC) is one of the ten most common cancers worldwide, with 573278 new cases and 212536 deaths reported globally in 2020. Currently, cystoscopy is the gold standard for diagnosing bladder cancer and remains the main method for assessing treatment effectiveness and monitoring recurrence. Due to the high recurrence rate of bladder cancer after treatment (up to 74.3% within 10 years), patients need to be followed up regularly (every 3-6 months) by cystoscopy for 5 years, followed by annual examination after tumor resection. This makes patient management in this disease the most expensive among all cancers, leading to increased economic burden, patient discomfort, and increased risk of complications. In contrast, non-invasive tests used in clinical diagnosis, such as urine cytology, are hindered from widespread adoption due to factors such as low sensitivity or specificity, susceptibility to various influences, high testing costs, etc. Therefore, to reduce the harm and economic burden on patients during testing, there is an urgent need to develop a non-invasive, convenient, and accurate diagnostic method.

[0229] For decades, urinary biomarkers have been a key area of bladder cancer research, spanning the experimental, clinical, and commercial fields. The concept is based on the continuous exposure of urine to the tumor surface, facilitating the non-invasive extraction of genomic, epigenetic, transcriptomic, and morphological information from shed cells. Unlike cell- and protein-based markers in urine, which are susceptible to conditions such as inflammation, urinary stones, hematuria, or epithelial dysplasia, tumor DNA in urine, which comes from the tumor, provides a valuable alternative to tissue-based genomic analysis.

[0230] With the advancement of next-generation sequencing (NGS) and digital polymerase chain reaction (dPCR) technologies, which help to detect tiny amounts of tumor DNA in a large amount of normal genomic DNA, the reliability and accessibility of urine DNA-based bladder cancer detection have been greatly improved. Compared with NGS, dPCR is more suitable for sensitive BC screening because of its significantly lower detection limit (0.001% sensitivity compared with 1% for NGS) and simpler process. However, relevant studies have shown that the mutation spectrum of bladder cancer is diverse, and there is a lack of universal hotspots. The incidence of the first three missense mutations ERBB2 S310F, PIK3CA E545K, and FGFR3 S249C is 4.7%, 6.6%, and 7.1%, respectively. Therefore, the practicability of the dPCR technology is limited by its limited multiplex detection capability. In 2013, two mutations were found in the Telomerase Reverse Transcriptase (TERT) promoter, g.1295228 C>T and g.1295250 C>T, which are prevalent in 59-77% of urothelial bladder cancer (UBC) cases. This finding provides a special opportunity for the development of a simple diagnostic test for BC using dPCR technology. Subsequent efforts to evaluate the clinical effectiveness of TERT promoter mutations (TERTpms) in BC identification showed a sensitivity of 52% to 87% and a specificity of 83% to 99%. In order to improve the detection sensitivity, combining multiple biomarkers in a diagnostic model has been proven to be effective in various disease studies. Given that abnormal DNA methylation is a common and early event in the process of bladder cancer, integrating TERT promoter mutations (TERTpms) and DNA methylation biomarkers into a diagnostic model is expected to be useful for detection. However, there is currently a lack of an effective means of high-sensitivity urine detection of bladder cancer.

[0231] The present scheme uses dPCR to detect low levels of TERT mutations and NRN1 methylation in urine with high sensitivity, and uses a machine learning model to optimize the detection weight of each biomarker, which can achieve accurate detection of patients with bladder cancer.

[0232] Please refer to FIG. 6, which shows a clinical sample experiment flowchart provided by an example embodiment of the present application. As shown in FIG. 6, the clinical sample experiment process is as follows:

[0233] 1) Urine sample collection and processing (i.e., urine collection)

[0234] For each enrolled subject, at least 20 ml of urine sample was collected from their first morning urine; then, these urine samples were separated into 10 ml of urine DNA storage tubes and stored at a temperature of 6-35 °C. Further processing was performed within 7 days.

[0235] 2) DNA extraction, sodium bisulfite treatment and methylation analysis

[0236] Genomic DNA and cell-free DNA (cfDNA) extracted from urine samples were performed using urine kits according to the corresponding manufacturer’s instructions. DNA quantification was performed by fluorometric quantitation.

[0237] That is, the urine sample was centrifuged at 3000 g for 15 minutes after adding Clearing Beads. Subsequently, the obtained precipitate was washed in digestion buffer and proteinase K, and then a DNA purification step was performed using a spin column.

[0238] Sodium bisulfite treatment was performed on 100-1000 ng of genomic DNA per sample using the DNA Methylation Sodium Bisulfite Kit according to the instructions with purified DNA as starting material. A total of 20 microliters of DNA was treated with 130 microliters of CT conversion mix in a thermocycler with the following conditions: 98 °C for 10 minutes, 64 °C for 40 minutes, 98 °C for 5 minutes, 64 °C for 40 minutes, 98 °C for 5 minutes, 64 °C for 40 minutes, and hold at 4 °C. The treated DNA was then purified using a DNA column.

[0239] 3) dPCR experiment and results output

[0240] For each dPCR assay, 10 microliters of reaction mix were prepared, including 1 microliter of template DNA, 6 microliters of 2x ddPCR Supermix, 2 microliters of forward and reverse primers (9 micromolar each), 1 microliter of FAM and HEX fluorescent probes for mutation detection (2.5 micromolar each for each mutation).

[0241] The reaction mix was introduced into the dPCR chip array microwells by injection and further introduced into each reaction microwell separated by dPCR oil. After sealing, the PCR reaction was performed under the following conditions: 96 °C for 10 minutes, 96 °C for 50 cycles, 30 seconds each cycle (slope 2.5 / sec), 55 °C for 30 seconds, then 60 °C for 10 minutes, and finally hold at 10 °C.

[0242] The primer and probe sequences of TERT promoter mutation are shown in Table 1 and Table 2, which are synthesized by Sheng Wu. The linear dynamic range detection process of TERT promoter mutation (C250T and C228T) is as follows:

[0243] According to the above dPCR reaction condition, the plasmid containing C250T and C228T synthesized by Sheng Wu is used as the DNA template substrate, and the dynamic linear range test of MAF of single C250T, C228T and single tube simultaneous detection of C250T and C228T is carried out. By adding different concentrations of mutant synthetic plasmid in wild type synthetic plasmid, the test range is from 0.5% to 100%, and the linear fitting result is shown in Figure 7, and the specific dPCR result is shown in Figure 8.

[0244] In Figure 8, part 81 is the dPCR result of the mutant synthetic plasmid, and part 82 is the dPCR result of the wild type synthetic plasmid. As can be seen from the result in Figure 7, the linear fitting results of single primer probe and mixing two kinds of mutant primer probe to the plasmid DNA template are very close. When the single tube simultaneous detection is carried out, the internal multiple primers and probes do not have great interference on the detection result of the dPCR system, so that the volume of dPCR reagent and sample required can be reduced by simultaneously detecting C250T and C228T, thereby the detection cost can be reduced.

[0245] The primer and probe sequences of NRN1 methylation are shown in Table 3, which are synthesized by Sheng Wu. The NRN1 detection is based on ATCB as the internal reference gene, and the linear range detection process of NRN1 methylation hotspot is as follows:

[0246] The NRN1 methylation hotspot chr6:6004463-6004464 is used as a detection marker, and the primer and probe sequences are shown in Table 3. The dPCR reaction condition is as described above, the test range is from 100 to 104 copies per microliter, and the linear fitting of the test result is shown in Figure 9. As can be seen from Figure 9, the primer and probe of the present embodiment can accurately quantitatively detect the content of the NRN1 methylation synthetic plasmid put into the experiment.

[0247] 4) Data analysis

[0248] The results of dPCR experiment, i.e. the mutation allele frequency MAF of TERT promoter mutation and the content of NRN1 methylation, are input into the classification model to obtain the classification result of the medical sample.

[0249] The Receiver Operating Characteristic Curve (ROC curve) is a graphical tool used to demonstrate the sensitivity and specificity of a binary classification system under different threshold settings. With the false negative rate (1-specificity) as the horizontal axis and the true positive rate (sensitivity) as the vertical axis, it is used to evaluate the ability of a model or test to distinguish between positive and negative classes. The Area Under the Curve (AUC) is an important indicator of classifier performance, ranging from 0.5 to 1, with a value closer to 1 indicating better classification performance.

[0250] That is, the performance of the model can be evaluated using the Area Under the Curve (AUC) statistic. The classifier in the AUC analysis consists of the methylation ratio of the biomarker or the minimum allele frequency (MAF) of the mutation in each sample. In mutation analysis, Pearson's chi-square test and Fisher's exact test were used to compare binary variables. The Shapiro-Wilk test was used to assess the normality of the data. Subsequently, for normally distributed and non-normally distributed data, Student's t-test and Mann-Whitney U-test were applied for comparison, respectively. The diagnostic test parameters including sensitivity, specificity, and accuracy were evaluated using the Clopper-Pearson method. In addition, the relationship between the methylation ratio or TERT promoter mutation MAF and other clinical variables (such as tumor stage and grade) was explored using the Mann-Whitney U-test.

[0251] The distribution of test results for different clinical samples (case and control groups) obtained by dPCR experiments is shown in Figure 10. Part (a) of Figure 10 is the dPCR test result distribution graph corresponding to TERT mutation, and part (b) of Figure 10 is the dPCR test result distribution graph corresponding to NRN1 methylation, where *** indicates p<0.001 and **** indicates p<0.0001. The p-value represents the statistical significance level, which is used to determine whether the observed results are unlikely to be caused by chance factors alone. As can be seen from the figure, there is a significant difference in the content of both TERT promoter mutation biomarkers and NRN1 methylation biomarkers between normal people and patients.

[0252] In addition, the above classification model can also be applied to the diagnosis of bladder cancer staging. The specific distribution results of the normal control group, the I group, and the II+ group are shown in FIG. 11. Part (a) of FIG. 11 is a dPCR test result distribution diagram corresponding to TERT mutation, and part (b) of FIG. 11 is a dPCR test result distribution diagram corresponding to NRN1 methylation. N.S. indicates no correlation, that is, there is not enough evidence to prove that there is a statistically significant relationship between variables. Among them, there is no significant difference between the I patients and the II and above patients from a single biological indicator. Medical personnel use a three-classification diagnosis model for analysis and prediction. Through comparison of different algorithms, it can be seen that the AUC value of the vector machine model (SMV) is the highest, and the specific results are shown in FIG. 12. Through diagnosis model prediction, normal people, I patients, and II+ patients can be distinguished to a certain extent, with an accuracy of 92.1%, 71.1%, and 78.9%, respectively. The results are shown in FIG. 13, indicating that it has a certain staging diagnosis capability. The ROC curve of the test set of the normal control group, the I patient, and the II+ patient model prediction is shown in part (a) of FIG. 13, and the test set confusion matrix diagram is shown in part (b) of FIG. 13.

[0253] In summary, the method shown in the above embodiments of the present application can be applied to a urine biomolecule multi-omics detection scheme. The scheme detects TERT mutations (C228T and C250T) and NRN1 methylation sites in urine, combines a machine learning model, realizes high-sensitivity detection of bladder cancer, and has a certain prediction effect on different stages of bladder cancer by optimizing the detection weight of the biomarker.

[0254] Please refer to FIG. 14, which shows a block diagram of a classification device for medical samples provided by an example embodiment of the present application. The device can be used to execute all or part of the steps performed by the computer device in the method shown in FIG. 2, as shown in FIG. 14. The device includes:

[0255] An information acquisition module 1401 is configured to acquire input information of a first medical sample. The input information of the first medical sample includes a gene measurement value of the first medical sample. The gene measurement value is used to indicate at least one of a gene mutation condition and a gene methylation condition.

[0256] An input module 1402 is configured to input the input information of the first medical sample into a classification model to obtain an output result of the classification model. The output result is used to indicate a probability distribution of the first medical sample in each candidate classification. The classification model is a machine learning model trained by input information of a second medical sample and a classification of the second medical sample.

[0257] A result acquisition module 1403 is configured to acquire a classification result of the first medical sample according to the output result. The classification result is used to indicate the classification of the first medical sample.

[0258] In some embodiments, the gene measurement value comprises at least one of a first gene measurement value and a second gene measurement value; the first gene measurement value is used to indicate a mutant allele frequency MAF of a TERT promoter mutation, and the second gene measurement value is used to indicate an NRN1 methylation content.

[0259] In some embodiments, the input information of the first medical sample further comprises one or more of the following information:

[0260] attribute information of a user corresponding to the first medical sample;

[0261] a detection result of abnormal cells in the first medical sample;

[0262] a medical image detection result of a tissue region corresponding to the first medical sample.

[0263] In some embodiments, the detection result of abnormal cells in the first medical sample comprises one or more of the following information:

[0264] a proportion of abnormal cells in the first medical sample;

[0265] a classification of abnormal cells in the first medical sample.

[0266] In some embodiments, the medical image detection result of the tissue region corresponding to the first medical sample comprises one or more of the following information:

[0267] a classification of the tissue region corresponding to the first medical sample;

[0268] an area proportion of an abnormal region in the tissue region corresponding to the first medical sample;

[0269] position information of the abnormal region of the tissue region corresponding to the first medical sample;

[0270] a medical image of the tissue region corresponding to the first medical sample.

[0271] In some embodiments, the first medical sample is at least one of the following medical samples:

[0272] a urine sample, a blood sample, and a tissue section sample.

[0273] In some embodiments, the input module is configured to multiply each of the plurality of data in the input information of the first medical sample by a respective weight of the data to obtain optimized input information.

[0274] The input module is configured to input the optimized input information into the classification model to obtain an output result of the classification model.

[0275] In some embodiments, each candidate classification includes a first candidate classification and a second candidate classification, the first candidate classification is a non-designated classification, and the second candidate classification is a designated classification; the classification of the second medical sample is one of the first candidate classification and the second candidate classification.

[0276] Or,

[0277] Each candidate classification includes a first candidate classification and at least two third candidate classifications, the first candidate classification is a non-designated classification, and the at least two third candidate classifications are at least two sub-classifications under the designated classification; the classification of the second medical sample is one of the first candidate classification and the at least two third candidate classifications.

[0278] In some embodiments, the apparatus further includes a training module configured to input the input information of the second medical sample into the classification model to obtain a prediction result output by the classification model; the prediction result is used to indicate a probability distribution of the second medical sample in each candidate classification.

[0279] The training module is configured to update model parameters of the classification model according to a difference between the probability distribution of the second medical sample in each candidate classification and the classification of the second medical sample.

[0280] In some embodiments, the classification model includes a plurality of network layers, and each of the plurality of network layers has a respective weight matrix.

[0281] The model parameters of the classification model include:

[0282] The respective weight matrix of each of the plurality of network layers.

[0283] Or,

[0284] The respective weight matrix of each of the plurality of network layers, and a respective weight of each of a plurality of data in the input information.

[0285] In some embodiments, the gene measurement value of the first medical sample is a gene measurement value obtained by performing dPCR detection on the first medical sample by using a kit.

[0286] In some embodiments, the kit includes a composition, and the composition includes a plurality of different decoy oligonucleotides, wherein the plurality of different decoy oligonucleotides are configured to collectively hybridize to a plurality of DNA molecules derived from a target genomic region.

[0287] Each of the target genomic regions is differentially methylated and / or differentially mutated in the at least one first sample compared to the second sample.

[0288] In some embodiments, the mutation refers to a TERT promoter region mutation, and the TERT promoter region mutation includes at least one of C228T and C250T.

[0289] In some embodiments, the plurality of DNA molecules is derived from at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic region of any one of SEQ ID NOs: 1-2.

[0290] In some embodiments, the plurality of DNA molecules is derived from at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic region of SEQ ID NOs: 1-2.

[0291] It should be noted that the apparatus provided by the above embodiments, when realizing its functions, is only exemplified by the above division of various functional modules, and in actual application, the above functions can be completed by different functional modules according to actual needs, that is, the content structure of the device is divided into different functional modules to complete all or part of the above described functions.

[0292] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method; the technical effects achieved by the operations of various modules are the same as those in the embodiments of the method, and will not be described in detail here.

[0293] Please refer to FIG. 15, which shows a structural block diagram of a computer device 1500 provided by an example embodiment of the present application. The computer device can be implemented as a server in the above-mentioned solutions of the present application. The computer device 1500 includes a central processing unit (CPU) 1501, a system memory 1504 including a random access memory (RAM) 1502 and a read-only memory (ROM) 1503, and a system bus 1505 connecting the system memory 1504 and the central processing unit 1501. The computer device 1500 also includes a mass storage device 1506 for storing an operating system 1509, application programs 1510, and other program modules 1511.

[0294] The mass storage device 1506 is connected to the central processing unit 1501 through a mass storage controller (not shown) connected to the system bus 1505. The mass storage device 1506 and its associated computer readable medium provide nonvolatile storage for the computer device 1500. That is, the mass storage device 1506 can include a computer readable medium (not shown) such as a hard drive or a compact disk read-only memory (CD-ROM) drive.

[0295] Without loss of generality, the computer readable medium can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes RAM, ROM, erasable programmable read only memory (EPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other solid state memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices. Of course, the computer storage media is not limited to the above-mentioned several kinds. The system memory 1504 and the mass storage device 1506 described above can be collectively referred to as memory.

[0296] According to various embodiments of the present disclosure, the computer device 1500 can also run on a remote computer connected to the network through a network such as the Internet. That is, the computer device 1500 can be connected to the network 1508 through the network interface unit 1507 connected to the system bus 1505, or can be connected to other types of networks or remote computer systems (not shown) using the network interface unit 1507.

[0297] The memory also includes at least one computer instruction stored in the memory, and the central processing unit 1501 implements all or part of the steps in the method shown in the above various embodiments by executing the at least one computer instruction.

[0298] In an exemplary embodiment, a chip is also provided, which includes programmable logic circuit and program instructions, when the chip runs on a computer device, for implementing the classification method of medical samples of the above aspects.

[0299] In an example embodiment, a computer program product is also provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor reads and executes the computer instructions from the computer readable storage medium to implement the classification method of the medical sample provided by each of the above method embodiments.

[0300] In an example embodiment, a computer readable storage medium is also provided, which stores computer instructions, and the computer instructions are loaded and executed by a processor to implement the classification method of the medical sample provided by each of the above method embodiments.

[0301] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0302] A person of ordinary skill in the art should be aware that, in one or more examples described above, the functions described in the embodiments of the present application can be implemented in hardware, software, firmware or any combination thereof. When implemented in software, the functions can be stored in a computer readable medium or transmitted as one or more instructions or code on a computer readable medium. The computer readable medium includes computer storage medium and communication medium, and the communication medium includes any medium that facilitates the transfer of computer program from one place to another. The storage medium can be any available medium accessible by a general or special purpose computer.

[0303] The above description is only optional embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method of classifying a medical sample, characterized by, The method comprises: obtaining input information of a first medical sample, the input information of the first medical sample comprising a gene measurement value of the first medical sample; the gene measurement value is used to indicate at least one of a gene mutation condition and a gene methylation condition; inputting the input information of the first medical sample into a classification model to obtain an output result of the classification model; the output result is used to indicate a probability distribution of the first medical sample in each candidate classification; the classification model is a machine learning model trained by input information of a second medical sample and a classification of the second medical sample; obtaining a classification result of the first medical sample according to the output result; the classification result is used to indicate the classification of the first medical sample.

2. The method of claim 1, wherein, The gene measurement value comprises at least one of a first gene measurement value and a second gene measurement value; the first gene measurement value is used to indicate a mutation allele frequency MAF of a telomerase reverse transcriptase TERT promoter mutation, and the second gene measurement value is used to indicate a neuromodulin NRN1 methylation content.

3. The method of claim 1, wherein, The input information of the first medical sample further comprises one or more of the following information: attribute information of a user corresponding to the first medical sample; a detection result of abnormal cells in the first medical sample; a medical image detection result of a tissue region corresponding to the first medical sample.

4. The method of claim 3, wherein, The detection result of abnormal cells in the first medical sample comprises one or more of the following information: a proportion of abnormal cells in the first medical sample; a classification of abnormal cells in the first medical sample.

5. The method of claim 3, wherein, The medical image detection result of the tissue region corresponding to the first medical sample comprises one or more of the following information: a classification of the tissue region corresponding to the first medical sample; an area proportion of an abnormal region in the tissue region corresponding to the first medical sample; position information of the abnormal region of the tissue region corresponding to the first medical sample; a medical image of the tissue region corresponding to the first medical sample.

6. The method of claim 1, wherein, The first medical sample is at least one of the following medical samples: a urine sample, a blood sample, and a tissue section sample.

7. The method of claim 1, wherein, The inputting of the input information of the first medical sample into the classification model to obtain the output result of the classification model comprises: multiplying a plurality of data in the input information of the first medical sample by respective weights of the plurality of data to obtain optimized input information; inputting the optimized input information into the classification model to obtain the output result of the classification model.

8. The method of claim 1, wherein the candidate classifications comprise a first candidate classification and a second candidate classification, the first candidate classification being a non-designated classification, and the second candidate classification being a designated classification; the classification of the second medical sample is one of the first candidate classification and the second candidate classification; or ​ The respective candidate classifications include a first candidate classification and at least two third candidate classifications, the first candidate classification being a non-designated classification, and the at least two third candidate classifications being at least two sub-classifications under a designated classification; the classification of the second medical sample is one of the first candidate classification and the at least two third candidate classifications.

9. The method according to any one of claims 1 to 8, characterized in that, Before the input information of the first medical sample is obtained, the method further includes: inputting the input information of the second medical sample into the classification model to obtain a prediction result output by the classification model; the prediction result is used to indicate a probability distribution of the second medical sample in the respective candidate classifications; updating model parameters of the classification model according to a difference between the probability distribution of the second medical sample in the respective candidate classifications and the classification of the second medical sample.

10. The method of claim 9, wherein, The classification model includes a plurality of network layers, and each of the plurality of network layers has a respective weight matrix; The model parameters of the classification model include: the weight matrix of each of the plurality of network layers; or, the weight matrix of each of the plurality of network layers, and a weight of each of a plurality of data in the input information.

11. The method according to any one of claims 1 to 10, characterized in that, The gene measurement value of the first medical sample is obtained by a kit through dPCR detection on the first medical sample; The kit includes a composition, and the composition includes a plurality of different decoy oligonucleotides, wherein the plurality of different decoy oligonucleotides are configured to collectively hybridize to a plurality of DNA molecules derived from a target genomic region; wherein each of the target genomic regions is differentially methylated and / or differentially mutated in at least one first sample compared to a second sample.

12. The method of claim 11, wherein, The mutation refers to a TERT promoter region mutation, and the TERT promoter region mutation includes at least one of C228T and C250T.

13. The method of claim 11, wherein, The plurality of DNA molecules are derived from at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic region in any one of sequences 1 to 2.

14. The method of claim 11, wherein, The plurality of DNA molecules are derived from at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic region in any one of sequences 1 to 2.

15. A composition characterized in that, The composition includes a plurality of different decoy oligonucleotides, wherein the plurality of different decoy oligonucleotides are configured to collectively hybridize to a plurality of DNA molecules derived from a target genomic region; wherein each of the target genomic regions is differentially methylated and / or differentially mutated in at least one first sample compared to a second sample.

16. The composition of claim 15, wherein, The mutation refers to a TERT promoter region mutation, and the TERT promoter region mutation includes at least one of C228T and C250T.

17. The composition of claim 15, wherein, The number of DNA molecules is at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic region derived from any one of SEQ ID NOs: 1-2.

18. The composition of claim 15, wherein, The number of DNA molecules is at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the target genomic region derived from SEQ ID NOs: 1-2.

19. A kit comprising, The kit comprises the composition of any one of claims 15-18.

20. A medical sample classification apparatus characterized by comprising: The device comprises: The first acquisition module is configured to acquire input information of a first medical sample, the input information of the first medical sample comprising a gene measurement value of the first medical sample, the gene measurement value being used to indicate at least one of a gene mutation condition and a gene methylation condition; The input module is configured to input the input information of the first medical sample into a classification model to obtain an output result of the classification model, the output result being used to indicate a probability distribution of the first medical sample in each candidate category, and the classification model being a machine learning model trained by input information of a second medical sample and a classification of the second medical sample; The second acquisition module is configured to acquire a classification result of the first medical sample according to the output result, the classification result being used to indicate the classification of the first medical sample. The computer device comprises a processor and a memory, and the memory stores at least one computer instruction, which is loaded and executed by the processor to implement the medical sample classification method according to any one of claims 1-14.

21. A computer device, comprising: The computer readable storage medium stores at least one computer instruction, which is loaded and executed by the processor to implement the medical sample classification method according to any one of claims 1-14.

22. A computer-readable storage medium, characterized in that, The computer program product comprises computer instructions stored in a computer readable storage medium, and the computer instructions are read and executed by a processor of a computer device to implement the medical sample classification method according to any one of claims 1-14.

23. A computer program product, characterised in that, ​