Gene marker and use thereof in risk prediction model for cervical cancer
Through quantitative detection and machine learning models of HM13, SNU13 and GNAS gene markers, the false positive and insufficient sensitivity of existing cervical cancer screening methods are solved, and early accurate detection of CIN3 and invasive cervical cancer is achieved, and the accuracy and sensitivity of screening are improved.
Patent Information
- Application Number
- PCT/CN2024/134933
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2024-11-27
- Publication Date
- 2025-07-10
AI Technical Summary
Existing cervical cancer screening methods such as HPV-DNA detection have false positive problems. HPV E6/E7 mRNA detection and cervical cell DNA methylation detection are insufficient in terms of sensitivity and cannot effectively detect precancerous cervical lesions and early cancers.
Quantitative detection at the single-cell level was performed using HM13, SNU13 and GNAS gene markers, combined with machine learning models, image information was obtained through quantitative chromogenic blot gene in situ hybridization technology, and characteristic features were extracted for cervical cancer risk prediction.
It improves the accuracy and sensitivity of cervical cancer screening, can detect CIN3 and invasive cervical cancer early, reduce unnecessary examinations and treatments, and saves medical resources.
Smart Images

Figure CN2024134933_10072025_PF_FP_ABST
Abstract
Description
Genetic markers and their application in cervical cancer risk prediction models Technical Field
[0001] The present application relates to the field of molecular biological detection, and specifically, to gene markers and their application in cervical cancer risk prediction models. Background Art
[0002] Cervical cancer, a malignant tumor with high morbidity and mortality among women, has been recognized as a major global public health problem. To address this challenge, in 2020, the World Health Organization (WHO) launched a global strategy aimed at accelerating the elimination of cervical cancer through increased HPV vaccination, high-precision screening, and appropriate treatment. Since over 90% of cervical cancers are caused by high-risk HPV, HPV DNA testing has become the most widely used cervical cancer screening method worldwide. However, because HPV can be cleared in most women without progressing to lesions, the specificity of HPV DNA testing is relatively low, which may lead to over-testing and over-treatment. To conserve medical resources, further triage of HPV-positive patients is necessary to reduce unnecessary colposcopies and biopsies. Currently, commonly used testing methods include HPV E6 / E7 mRNA testing and cervical DNA methylation testing. However, clinical application results show that HPV E6 / E7 mRNA testing does not significantly improve specificity compared to HPV DNA testing and still has a high false positive rate. Although cervical cell DNA methylation testing can detect invasive cervical cancer relatively accurately, its sensitivity for high-grade cervical intraepithelial neoplasia (CIN3, including cervical precancerous lesions and carcinoma in situ) is less than 70%, which cannot meet the current diagnostic and treatment guidelines for early treatment of CIN3.
[0003] Therefore, developing a new detection technology to accurately detect CIN3 and invasive cervical cancer has become an urgent clinical need. Summary of the Invention
[0004] This application is completed by the inventor based on the following findings:
[0005] Currently, cervical cancer screening mainly relies on HPV-DNA testing. However, this method has a serious false positive problem. To improve the accuracy of screening, scientists have introduced auxiliary detection methods, such as HPV E6 / E7 mRNA detection and cervical cell DNA methylation detection.
[0006] HPV E6 / E7 mRNAs are two products of human papillomavirus expression, and these two proteins play a key role in the carcinogenesis of cervical cells. However, in most women, E6 / E7 mRNA expression levels within a short period of time after HPV infection are insufficient to cause cell carcinogenesis. Therefore, HPV E6 / E7 mRNA testing still has similar false positive issues as HPV DNA testing.
[0007] On the other hand, changes in methylation of cervical cell DNA are an early event in cervical cancer. Current DNA methylation detection techniques require sulfite conversion of DNA samples to convert unmethylated cytosine bases (C) to uracil bases (U). However, this conversion process results in significant DNA loss, reducing sensitivity to methylation. Therefore, in practical applications, it has low sensitivity for detecting early lesions such as cervical precancerous lesions and carcinoma in situ.
[0008] In response to these challenges, there is an urgent need for a more accurate and sensitive cervical cancer screening method that can detect potential signs of cancer at an early stage and provide patients with earlier intervention and treatment opportunities. To this end, one object of the present invention is to provide a method that can effectively detect CIN3 and invasive cervical cancer.
[0009] In light of this, in the first aspect of this application, a genetic marker is proposed. According to an embodiment of this application, the genetic marker includes at least one of the following genes: HM13 (Minor Histocompatibility Antigen H13) and SNU13 (Small Nuclear Ribonucleoprotein 13). According to the embodiments of this application, the inventors have for the first time selected the expression changes of HM13 and SNU13, the genetic markers that occur at the earliest stages of cancer, as targets for cervical cancer detection. This quantitative detection is performed at the single-cell level, enabling early and accurate detection of CIN3 and invasive cervical cancer while excluding other low-risk lesions and avoiding unnecessary colposcopy and biopsy.
[0010] According to an embodiment of the present application, the above-mentioned gene markers may further include at least one of the following technical features:
[0011] According to embodiments of the present application, the gene marker further includes the GNAS (Guanine Nucleotide-binding Protein, Alpha-stimulating Complex Locus) gene. In some examples of the present application, combining HM13, SNU13, and GNAS genes for the detection of CIN3 and invasive cervical cancer can increase the sensitivity and accuracy of detection.
[0012] In one example of the present application, the gene marker further includes SNRPN (Small Nuclear Ribonucleoprotein Polypeptide N).
[0013] In the second aspect of this application, the present application proposes the use of a reagent for detecting the gene markers described in the first aspect in the preparation of a kit. According to embodiments of this application, the reagent is used to diagnose cervical intraepithelial neoplasia (CIN3) and invasive cervical cancer. In some examples of this application, by preparing reagents for detecting HM13 and SNU13 / HM13, or SNU13 and GNAS genes into a kit, cervical cancer detection is performed. In practical applications, the procedure is simpler and more cost-effective, making it more suitable for widespread cervical cancer screening.
[0014] According to an embodiment of the present application, the above-mentioned use may further include at least one of the following technical features:
[0015] According to embodiments of the present application, the reagent includes at least one of a probe, a primer, and a mass spectrometry detection reagent specific for the gene marker. In some examples of the present application, the reagent can specifically and highly sensitively detect the HM13, SNU13, and GNAS genes and screen samples for CIN3 and invasive cervical cancer.
[0016] According to the embodiments of the present application, the probes have nucleotide sequences shown in SEQ ID NOs: 1 to 144. Specific sequences are shown in Tables 1 to 4.
[0017] In a third aspect of this application, a kit for screening biological samples for cervical intraepithelial neoplasia (CIN3) and invasive cervical cancer is provided. According to embodiments of this application, the kit comprises a reagent suitable for detecting at least one of the gene markers described in the first aspect. In some examples of this application, the kit has higher accuracy and sensitivity for detecting CIN3 and invasive cervical cancer, and has advantages such as portability and wide applicability.
[0018] According to an embodiment of the present application, the above-mentioned kit may further include at least one of the following technical features:
[0019] According to embodiments of the present application, the reagents include at least one of a probe, primer, and mass spectrometry detection reagent specific for the gene marker. In some examples of the present application, the use of the above-mentioned reagents to detect HM13, SNU13, and GNAS genes is simple to operate and cost-effective in practical applications, and is widely used in the screening and prevention of cervical cancer.
[0020] According to an embodiment of the present application, the reagent includes a probe specific for the gene marker, and the probe has a nucleotide sequence shown in SEQ ID NOs: 1 to 144. Specific sequences are shown in Tables 1 to 4.
[0021] According to an embodiment of the present application, the kit further comprises at least one of a fluorescent reagent, a buffer solution and a washing solution. In some examples of the present application, the probe is typically combined with a fluorescent reagent or other detection labels to detect the expression status of the gene in the aforementioned marker.
[0022] In a fourth aspect, this application provides a drug for treating cervical intraepithelial neoplasia and invasive cervical cancer. According to embodiments of this application, the drug comprises an agent that specifically alters the gene markers described in the first aspect. In some examples of this application, the drug is used to treat cervical intraepithelial neoplasia and invasive cervical cancer.
[0023] It should be noted that the so-called "specific changes" refer to the ability to restore the mutated nucleic acid or mutated protein to its original wild state or other non-pathogenic state without having a substantial impact on other sequences of the individual's genome.
[0024] According to an embodiment of the present application, the above-mentioned medicine may further include at least one of the following technical features:
[0025] According to an embodiment of the present application, the drug further includes: a pharmaceutically acceptable carrier or excipient.
[0026] As used herein, the term "pharmaceutically acceptable" indicates that the pharmaceutical composition can be administered to a subject without causing adverse physiological reactions that would prevent the administration of the pharmaceutical composition.
[0027] The term "pharmaceutically acceptable carrier" may include any and all physiologically compatible solvents. Specific examples include one or more of water, saline, phosphate-buffered saline, glucose, and combinations thereof. In many cases, pharmaceutical compositions include isotonic agents, such as sodium chloride. Of course, pharmaceutically acceptable carriers may also include trace amounts of auxiliary substances, such as buffers, to extend the shelf life or potency of the antibody.
[0028] The "pharmaceutically acceptable excipients" are excipients useful in preparing generally safe, non-toxic and desirable pharmaceutical compositions. Preferably, examples of these excipients or diluents include, but are not limited to, water, saline, Ringer's solution, glucose, mannitol, dextrose, lactose, starch, magnesium stearate, cellulose, magnesium carbonate, 0.3% glycerol, hyaluronic acid, ethanol, polyalkylene glycols such as polypropylene glycol, triglycerides, and 5% human serum albumin. Liposomes and non-aqueous vehicles, such as fixed oils, may also be used.
[0029] In a fifth aspect, the present application provides an intron probe. According to embodiments of the present application, the probe is used to detect the gene markers described in the first aspect. In some examples of the present application, the intron probe can be used to specifically detect at least one of the gene markers HM13, SNU13, and GNAS.
[0030] According to an embodiment of the present application, the above-mentioned intron probe may further include at least one of the following technical features:
[0031] According to the embodiments of the present application, the probes have nucleotide sequences shown in SEQ ID NOs: 1 to 144. Specific sequences are shown in Tables 1 to 4.
[0032] In a sixth aspect, this application proposes a method for training a machine learning model for predicting the risk of cervical intraepithelial neoplasia and invasive cervical cancer. According to an embodiment of this application, the method comprises: obtaining in situ hybridization image information of introns of gene markers in a training sample, the image information comprising hybridization signals obtained by hybridizing the gene markers described in the first aspect with the probes described in the fifth aspect; obtaining predetermined features of the training sample based on the image information; inputting the predetermined features into a machine learning model, and using the known risk status of the training sample as a marker to perform supervised training on the machine learning model to obtain the machine learning model. In some examples of this application, this method achieves highly accurate and personalized prediction of individual abnormality status by combining high-resolution image information with machine learning technology. The inventors perform in situ hybridization of gene markers in cervical samples using specific probe pairs to detect the allelic expression status of the gene markers in the cell nucleus. The abnormal expression status of the gene markers is further quantitatively analyzed. The results of the quantitative analysis are converted into a cervical cancer risk index by combining the machine learning model. Judging the benign or malignant nature of samples based on the risk index provides an important reference for the early diagnosis of cervical cancer. This method provides a highly accurate and reliable means of cervical cancer screening, providing clinicians with a more accurate diagnosis and offering patients the opportunity to receive treatment and intervention earlier.
[0033] According to an embodiment of the present application, the above-mentioned machine learning model training method may also have at least one of the following additional technical features:
[0034] According to an embodiment of the present application, the machine learning model is selected from at least one of logistic regression, neural network, decision tree and random forest. In some preferred examples of the present application, the machine learning model is selected from logistic regression.
[0035] In some examples of this application, the machine learning model training method described herein can be used not only to train prediction models for the risk of cervical intraepithelial neoplasia and invasive cervical cancer, but also for prediction models for lung cancer, thyroid cancer, breast cancer, and bladder cancer. Accordingly, the image information is selected from image information of corresponding gene markers.
[0036] In one example of this application, the inventors first used quantitative chromogenic in situ hybridization (QCIGISH) to obtain intronic in situ hybridization images of the gene markers GNAS, HM13, and SNU13. Using this image information, they extracted key features and combined them with a machine learning model to predict cervical cancer risk. After comprehensively considering model computation speed and prediction accuracy, the inventors ultimately discovered that extracting features from the image information of the gene markers GNAS, HM13, and SNU13 yielded superior model performance.
[0037] According to an embodiment of the present application, the predetermined characteristics include at least one selected from the group consisting of the proportion of cells with two hybridization signals to all cells (BAE), the proportion of cells with three and four hybridization signals to all cells (MAE3-4), the proportion of cells with five and six hybridization signals to all cells (MAE5-6), the proportion of cells with seven and eight hybridization signals to all cells (MAE7-8), the proportion of cells with nine or more hybridization signals to all cells (MAE9+), and the proportion of cells with hybridization signals to all cells (TE).
[0038] In some preferred examples of the present application, the predetermined feature includes at least one selected from the group consisting of the proportion of cells with two hybridization signals among all cells, the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, the proportion of cells with seven and eight hybridization signals among all cells, and the proportion of cells with no hybridization signals among all cells. The inventors optimized the features and deleted some less important features to improve model performance, reduce computational costs, and increase model interpretability.
[0039] In some more preferred examples of the present application, the inventors further optimized model performance by combining gene markers with the aforementioned features to adapt the optimal characteristics of each gene marker. In some preferred models, the predetermined characteristics of the gene markers GNAS and SNU13 are selected from at least one of the following: the proportion of cells with two hybridization signals among all cells, the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, and the proportion of cells with no hybridization signals among all cells; and the predetermined characteristics of the gene marker HM13 are selected from at least one of the following: the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, the proportion of cells with seven and eight hybridization signals among all cells, and the proportion of cells with no hybridization signals among all cells.
[0040] In the present application, there is no specific limitation on the method for obtaining the image. In some preferred examples of the present application, the image information is obtained by quantitative chromogenic imprinted gene in situ hybridization (QCIGISH) method.
[0041] In the seventh aspect of the present application, the present application proposes a cervical cancer risk prediction model. According to an embodiment of the present application, the prediction model includes: an image information acquisition module, the image information acquisition module is used to obtain the gene marker image information of the first aspect of the sample to be tested; a predetermined feature acquisition module is used to obtain the predetermined features of the training sample based on the image information; and a prediction module, the prediction module is used to input the predetermined features into a pre-trained machine learning model to obtain a prediction result. The pre-trained machine learning model is obtained by training using the method described in the sixth aspect. In some examples of the present application, the prediction model is used to predict the risk of cervical cancer.
[0042] According to an embodiment of the present application, the above prediction model may further include at least one of the following technical features:
[0043] In some examples of the present application, the predetermined features input into the pre-trained machine learning model may further include text features. In some examples of the present application, the text features include at least one of age information, body mass index, and medical history of the sample to be tested.
[0044] According to an embodiment of the present application, the sample to be tested is selected from cervical cell tissue.
[0045] According to an embodiment of the present application, the prediction result is determined as follows: if the output result of the machine learning model is greater than or equal to 0.466, the sample is judged to be a cervical cancer positive sample; if the output result of the machine learning model is less than 0.466, the sample is judged to be a cervical cancer positive sample.
[0046] It should be noted that, in this application, a positive cervical cancer sample means that the sample is predicted to be positive for CIN3 and invasive cervical cancer.
[0047] It should be noted that the above thresholds are obtained by the inventors of this application based on the machine learning model of this application and the number of training samples. In some other embodiments, model switching or different numbers of samples may cause changes in the thresholds.
[0048] In the eighth aspect of the present application, the present application proposes a method for predicting or diagnosing the risk of cervical cancer. According to an embodiment of the present application, the method includes: obtaining in situ hybridization image information of introns of the gene marker of the sample to be tested, the image information contains a hybridization signal, and the hybridization signal is obtained by hybridizing the gene marker of the first aspect with the probe described in the fifth aspect of the present application; based on the image information, obtaining a predetermined feature of the sample to be tested; inputting the predetermined feature into a pre-trained machine learning model to obtain a prediction result. Wherein, the pre-trained machine learning model is obtained by training the method described in the sixth aspect. In some examples of the present application, the method is used to predict or diagnose the risk of cervical cancer.
[0049] According to an embodiment of the present application, the above method may further include at least one of the following technical features:
[0050] In some examples of the present application, the predetermined features input into the pre-trained machine learning model may further include text features. In some examples of the present application, the text features include at least one of age information, body mass index, and medical history of the sample to be tested.
[0051] According to an embodiment of the present application, the sample to be tested is selected from cervical cell tissue.
[0052] According to an embodiment of the present application, the prediction result is determined as follows: if the output result of the machine learning model is greater than or equal to a predetermined threshold, the sample is judged to be a cervical cancer positive sample; if the output result of the machine learning model is less than the predetermined threshold, the sample is judged to be a cervical cancer positive sample.
[0053] According to an embodiment of the present application, the predetermined threshold is selected from 0.4-0.5, optionally 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49 or 0.5. According to a preferred embodiment of the present application, the predetermined threshold is selected from 0.466.
[0054] It should be noted that, in this application, a positive cervical cancer sample means that the sample is predicted to be positive for CIN3 and invasive cervical cancer.
[0055] It should be noted that the above thresholds are obtained by the inventors of this application based on the machine learning model of this application and the number of training samples. In some other embodiments, model switching or different numbers of samples may cause changes in the thresholds.
[0056] In some examples of the present application, after the sample is determined to be negative or positive by the above-mentioned determination method, further examination and treatment including medication and surgery can be performed on the cervical lesions. For example, negative cases can undergo liquid-based cytology examination every three months, or interferon antiviral treatment. For positive cases, further colposcopy and colposcopic biopsy are required. If it is determined to be CIN3, ablative treatments such as laser, electrocautery, and cryoablation can be performed, or cervical conization can be performed; if it is determined to be invasive cervical cancer, cervical conization or radical hysterectomy is required, and radiotherapy or chemotherapy can be performed as needed to eliminate metastatic cancer cells.
[0057] In the ninth aspect of the present application, the present application proposes a machine learning model training device. According to an embodiment of the present application, the device includes: an image acquisition unit for acquiring in situ hybridization image information of introns of gene markers in a training sample, wherein the image information includes a hybridization signal, and the hybridization signal is obtained by hybridizing the gene markers described in the first aspect with the probe described in the fifth aspect; a feature acquisition unit for acquiring predetermined features of the training sample based on the image information; and a training unit for inputting the predetermined features into a machine learning model, using the known risk status of the training sample as a marker, so as to perform supervised training on the machine learning model, so as to obtain the machine learning model.
[0058] It should be noted that the machine learning model training device described in this application is an extended application of the machine learning model training method described in the first aspect of this application. Therefore, the characteristics or advantages of the machine learning model training method described in the first aspect of this application are also applicable to this aspect and will not be repeated here.
[0059] In some examples of the present application, the device flow chart is shown in FIG1 , wherein the image acquisition unit S100 is connected to the feature acquisition unit S200 , and the feature acquisition unit S200 is connected to the training unit S300 .
[0060] In the tenth aspect of the present application, the present application proposes a server, which includes a processor and a memory, and a computer program is stored in the memory. When the computer program is executed by the processor, the machine learning model training method described in the sixth aspect is implemented.
[0061] In the eleventh aspect of the present application, the present application proposes a computer-readable storage medium containing a computer program, characterized in that when the computer program is executed by one or more processors, the machine learning model training method described in the sixth aspect is implemented.
[0062] It should be noted that the various embodiments of the methods and models described above can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0063] The program code for implementing the method disclosed in the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0064] In the context disclosed herein, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0065] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0066] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0067] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0068] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0069] It should be noted that the features and technical effects described in this article for different aspects can be used as reference for each other and will not be repeated here.
[0070] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0072] FIG1 is a flow chart of a machine learning model training apparatus according to an embodiment of the present application;
[0073] FIG2 is a schematic diagram of cell nuclei expressing different numbers of alleles observed under a bright field microscope according to one embodiment of the present application;
[0074] FIG3 is a schematic diagram of the detection results of the allele expression status of the HM13 gene observed under a bright field microscope according to one embodiment of the present application;
[0075] FIG4 is a schematic diagram of the detection results of the allele expression status of the SNU13 gene observed under a bright field microscope according to one embodiment of the present application;
[0076] FIG5 is a schematic diagram of the detection results of the allele expression status of the GNAS gene observed under a bright field microscope according to one embodiment of the present application;
[0077] FIG6 is a schematic diagram of the detection results of the allele expression status of the SNRPN gene observed under a bright field microscope according to one embodiment of the present application;
[0078] FIG7 is a schematic diagram of BAE, MAE and TE parameter analysis results of low-risk group and high-risk group samples according to one embodiment of the present application;
[0079] FIG8 is a schematic diagram of ROC curve results of BAE, MAE and TE parameters of GNAS, SNRPN, HM13 and SNU13 genes according to one embodiment of the present application;
[0080] FIG9 is a schematic diagram of cell nuclei expressing different numbers of alleles observed under a fluorescence microscope according to one embodiment of the present application;
[0081] FIG10 is a schematic diagram of the detection results of the allele expression status of the HM13 gene observed under a fluorescence microscope according to one embodiment of the present application;
[0082] FIG11 is a schematic diagram of the detection results of the allele expression status of the SNU13 gene observed under a fluorescence microscope according to one embodiment of the present application;
[0083] FIG12 is a schematic diagram of the detection results of the allele expression status of the GNAS gene observed under a fluorescence microscope according to one embodiment of the present application. DETAILED DESCRIPTION
[0084] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0085] As used herein, unless otherwise indicated, the singular forms "a," "an," and the like include plural referents (more than one); "a set" or "a plurality" refers to two or more.
[0086] In this document, unless otherwise specified, the terms "first", "second", "third", "fourth", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated; features specified as "first", "second", etc. may explicitly or implicitly include one or more of the said features.
[0087] Unless otherwise specified in this article, the terms "gene markers" and "imprinted genes" can be used interchangeably. As a special type of gene that is regulated by epigenetics, it comes from two alleles from the father and the mother. Under normal circumstances, only one is expressed, and the other is highly methylated and silent. In the earliest stage of tumor development, one of the originally silent alleles of the imprinted gene becomes demethylated and begins to be expressed again. This phenomenon is called loss of imprinting (LOI). LOI of imprinted genes is commonly found in many cancers, but there are no research results with clinical application value in cervical cancer. In this application, the inventors discovered three imprinted genes (GNAS, HM13 and SNU13) that can be used to assess the risk of cervical cancer. By obtaining the QCIGISH image information of the aforementioned imprinted genes and combining it with a machine learning model, accurate prediction of the risk of cervical cancer is achieved.
[0088] Unless otherwise specified herein, the term "Quantitative Chromogenic In Situ Hybridization (QCIGISH)" is a method for detecting the presence, quantity and status of certain specific genes or chromosome regions in cells or tissues. This method combines in situ hybridization (ISH) and chromogenic blot (Chromogenic Detection) technology, allowing direct positioning and detection of specific DNA or RNA sequences in cells or tissues. This technology is commonly used in cancer research and diagnosis to determine the presence, absence or quantitative changes of certain cancer-related genes in cells. In this application, the RNA in the sample is bound to specific probes, which are labeled with specific chromogenic markers to make the target sequence visible through a staining reaction. The presence and / or quantity of the target sequence is quantified by the intensity of the chromogenic reaction or the area of staining.
[0089] Unless otherwise specified, the terms "predictive model" and "machine learning model" in this document refer to a computational model or algorithm that automatically performs tasks such as prediction, classification, recognition, or decision-making by learning and analyzing input data. The model's learning process is based on statistical principles and data pattern recognition, using a training set to adjust model parameters and optimize the model to improve its predictive or inference capabilities. Machine learning models can employ various algorithms and techniques, such as neural networks, support vector machines, decision trees, random forests, and deep learning. These models can be trained and optimized through supervised learning, unsupervised learning, or reinforcement learning. In practical applications, machine learning models can be used in a variety of fields, such as natural language processing, image recognition, pattern recognition, data mining, recommender systems, and predictive analytics. They hold significant potential for large-scale data processing, automated decision-making, and intelligent systems. However, it should be noted that in the specific application of machine learning models, in-depth research on the features and models used for prediction is required to achieve satisfactory prediction results. Otherwise, various problems may arise, such as overfitting, underfitting, and poor generalization. After years of research and combined with research experience, the inventors of this application unexpectedly discovered a predetermined feature combination that can be combined with a machine learning model to assess the risk of cervical cancer.
[0090] In this article, supervised learning is used to train the machine learning model. So-called "supervised learning" is one of the most common and widely used learning methods in machine learning. The model is trained using training data with correct answers (labels), allowing the model to learn the relationship between input and output from the data, thereby making predictions or classifications on unseen data. In supervised learning, each training sample consists of input data and a corresponding label. The label in the training sample indicates what output should be produced for the given input data. The goal of the model is to learn a function from the training data that can map the input data to the corresponding output label. This process can be seen as learning the mapping relationship between input and output, also known as the model's "hypothesis function" or "mapping function."
[0091] In this article, we use supervised learning to solve classification problems. The goal of a machine learning model is to map input data to one of predefined discrete categories. The training data consists of a set of input examples and their category labels. After training, the model is able to classify new input examples and assign them to the correct category.
[0092] The training of machine learning models requires a training set and a test set. The so-called "training set" is a data set used to train the machine learning model. It contains a series of input samples (feature vectors) and corresponding known outputs (labels or target values). During the training phase, the machine learning model uses the samples in the training set to learn and adjust the parameters of the model to minimize the error between the predicted results and the true labels. By constantly trying different parameters and algorithms, the model gradually learns the patterns and characteristics between samples, thereby obtaining more accurate prediction capabilities. In this application, the training set comes from 75 cervical cell scraping samples from West China Second Hospital of Sichuan University.
[0093] The so-called "test set" is a data set used to evaluate the performance of a machine learning model. It also contains input samples and corresponding true outputs (labels), but the samples in the test set have not been used in the training phase of the model, which ensures that the data in the test set is unknown to the model. After the training is completed, the test set is used to evaluate the performance of the model, that is, the model predicts the samples in the test set and compares them with the true labels. By comparing the predicted results and the true labels, the performance of the model on unknown data is evaluated to determine its generalization ability and prediction accuracy. In this application, the test set comes from 105 cervical cell scraping samples from West China Second Hospital of Sichuan University.
[0094] Unless otherwise specified, the term "exhaustive approach" refers to exhaustively trying different coefficients to identify the optimal combination to combine the prediction results of genes. In this example, the inventors attempted to combine the prediction results of individual imprinted genes using different coefficients to identify the optimal combination model. By systematically trying different coefficient combinations and then evaluating the accuracy of each combination, they found the most effective coefficient combination for cervical cancer risk prediction.
[0095] Those skilled in the art will understand that during gene expression, DNA is transcribed into a molecule called nascent RNA. This nascent RNA contains the intron sequence of the gene. However, during transcription or shortly after transcription is completed, RNA undergoes a splicing process, intron sequences are cut out and rapidly degraded. Therefore, RNA containing intron sequences only exists near its transcription site. The inventors innovatively came up with the idea of using probes that recognize intron sequences of imprinted genes (gene markers as described above) for RNA in situ hybridization, which can mark the spatial location of gene transcription in the cell nucleus and form a signal visible under a microscope. By capturing images of the results of in situ hybridization, cells can be divided into four types according to the number of imprinted gene expression sites in the cell nucleus: no imprinted gene expression, one imprinted gene expression, two imprinted gene expression, and three or more imprinted gene expression; among them, cells expressing three or more imprinted genes can be further divided into four types: three or four imprinted gene expression, five or six imprinted gene expression, seven or eight imprinted gene expression, and nine or more imprinted gene expression.
[0096] It should be noted that in normal single cells, imprinted genes are not expressed or only one imprinted gene is expressed.
[0097] In this application, the inventors screened imprinted genes and trained and evaluated machine learning models based on the obtained imprinted genes, ultimately achieving accurate prediction of cervical cancer risk scores.
[0098] For ease of understanding, the following describes in detail the technical solution for applying the gene markers of this application in the risk prediction model for predicting CIN3 and invasive cervical cancer.
[0099] 1. Imprinted gene screening and specific probe design
[0100] After extensive literature analysis and research experience, the inventor unexpectedly discovered that imprinted genes can be used to assess cancer risk, and then for the first time proposed that at least one of the imprinted genes GNAS, HM13, SNRPN and SNU13 be used for cervical cancer risk prediction.
[0101] First, cervical cancer tissue samples are collected and tissue sections are prepared. The cervical cancer tissue samples include samples with varying degrees of lesions. Furthermore, the tissue sections are prepared using conventional methods in the art. In one example of the present application, tissue sections are prepared using the formalin-fixed paraffin-embedded (FFPE) method.
[0102] Secondly, with the specific intron of imprinted gene as target, design specific probe.In the present application, described intron probe design is not limited to certain or some specific introns, on the basis of being able to meet the above-mentioned design principle, the selection of intron is arbitrary, can select a specific intron (such as the 8th intron of imprinted gene HM13, etc.), also can select multiple introns to combine (such as the 8th, 10th, 15th intron combination of imprinted gene HM13).Each imprinted gene designs at least 3 pairs of specific intron probes.In some examples of the present application, each imprinted gene optionally designs 5,8,10,12,15,16,18,20,25 or 30 pairs of specific intron probes.In some preferred examples of the present application, each imprinted gene designs 15 pairs of specific intron probes, more preferably 16 pairs.
[0103] Among them, the design principles of specific probes should meet at least one of the following conditions: 1) at least three pairs of probes, each probe has a sequence length of 18-25 nucleotides that is complementary to the intron of the target gene; 2) under in situ hybridization conditions, the Tm value of each probe is equal to or higher than 60°C; 3) the distance between the binding sites of the two probes (first probe and second probe) of the same pair and the target sequence is less than or equal to 2 nucleotides; 4) any first probe and any second probe of the same gene cannot simultaneously have a continuous sequence of 15 nucleotides or longer that is continuously complementary to the introns of naturally occurring mRNA, non-coding RNA and other genes in the human body, among which non-coding RNA includes long non-coding RNA (lncRNA), micro RNA (miRNA), ribosomal RNA (rRNA), and transfer RNA (tRNA). Continuous complementary pairing means that the distance between the binding sites of the two probes on the same nucleotide is less than or equal to 2 nucleotides.
[0104] In a specific example of this application, based on the aforementioned specific probe design method, the inventors designed intron-specific probes for the imprinted genes GNAS, HM13, SNRPN, and SNU13 for RNA in situ hybridization. These probes have the nucleotide sequences set forth in SEQ ID NOs: 1-144. Specific sequences are shown in Tables 1-4.
[0105] Furthermore, based on the above method, a specific intron probe combination of imprinted genes was obtained. The QCIGISH method was used to perform in situ hybridization and chromogenic imprinting signal amplification, obtain the result image, and record the expression status of the imprinted gene in the cell nucleus. According to the counting results, the following formula was used to calculate the expression parameter of each imprinted gene: Imprinted gene biallelic expression (BAE) = N2 / (N1+N2+N3+ )×100%; Multiallelic expression of imprinted genes (MAE)=N 3+ / (N1+N2+N 3+ )×100%; Total expression of imprinted genes (TE)=(N1+N2+N 3+ ) / (N0+N1+N2+N 3+ )×100%;
[0106] Among them, N0 is the number of cell nuclei without HM13 gene expression sites, N1 is the number of cell nuclei with one HM13 gene expression site, N2 is the number of cell nuclei with two HM13 gene expression sites, and N 3+ is the number of nuclei with three or more HM13 gene expression sites.
[0107] Based on these imprinted gene expression parameters, we analyzed the expression of each gene in samples with varying disease severity and the association between gene expression and different clinical states. We calculated the BAE, MAE, and TE parameters for each imprinted gene and performed receiver operating characteristic (ROC) curve analysis. Imprinted gene screening was performed based on the area under the ROC curve.
[0108] In a specific example of the present application, the imprinted genes GNAS, HM13 and SNU13 are preferably used in combination for cervical cancer risk prediction.
[0109] 2. Model training and evaluation
[0110] Based on the imprinted genes obtained through the above screening and the specific intron probe combinations corresponding to each imprinted gene, the QCIGISH method was used to obtain the expression parameters of each imprinted gene, which were the input features of the machine learning model.
[0111] In some examples of the present application, the machine learning model is selected from at least one of logistic regression, neural network, decision tree, and random forest.
[0112] In some examples of the present application, the input features of the machine learning model are selected from at least one of the proportion of cells with two hybridization signals among all cells (BAE), the proportion of cells with three and four hybridization signals among all cells (MAE3-4), the proportion of cells with five and six hybridization signals among all cells (MAE5-6), the proportion of cells with seven and eight hybridization signals among all cells (MAE7-8), the proportion of cells with more than nine hybridization signals among all cells (MAE9+), and the proportion of cells with hybridization signals among all cells (TE).
[0113] To further improve model performance and prediction accuracy, the inventors optimized the input features of the machine learning model and selected the optimal feature parameters for each imprinted gene. The optimal feature parameters were selected based on the discrimination of the parameter combinations using the rfe and rfeControl functions in the R package caret.
[0114] In one example of the present application, the characteristic parameters of the imprinted genes GNAS and SNU13 are selected from at least one of the following: the proportion of cells with two hybridization signals among all cells, the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, and the proportion of cells with no hybridization signals among all cells; the characteristic parameters of the imprinted gene HM13 are selected from at least one of the following: the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, the proportion of cells with seven and eight hybridization signals among all cells, and the proportion of cells with no hybridization signals among all cells.
[0115] Based on the optimal characteristic parameters of each imprinted gene obtained above, model construction is performed for each imprinted gene. In a preferred example of the present application, the model is selected from a decision tree, and the area under the ROC curve of each gene modeled separately and the sensitivity and specificity in the training set are obtained. Based on the obtained area under the ROC curve, sensitivity and specificity results, the prediction results of two or three imprinted genes are combined, and an exhaustive method is used to select the best combination coefficient, and the accuracy of different combinations is evaluated to finally obtain the optimal prediction model. In some examples of the present application, the combination coefficient may vary based on the selection of the model and the number of training samples. Therefore, the specific coefficients and thresholds are not specifically defined in the present application.
[0116] Based on the optimal characteristic parameters of each imprinted gene obtained above, model construction is performed for each imprinted gene. In a preferred example of the present application, the model is selected from logistic regression, and the area under the ROC curve of each gene modeled separately and the sensitivity and specificity in the training set are obtained. Based on the obtained area under the ROC curve, sensitivity and specificity results, the prediction results of two or three imprinted genes are combined, and an exhaustive method is used to select the best combination coefficient, and the accuracy of different combinations is evaluated to finally obtain the optimal prediction model. In some examples of the present application, the combination coefficient may vary based on the choice of model and the number of training samples. Therefore, the specific coefficients and thresholds are not specifically defined in the present application.
[0117] 3. Model Validation
[0118] Obtain validation set samples (with no overlap with training set samples) containing known cervical cancer diagnosis results. Obtain model input features for the validation set samples using the above-described method for calculating imprinted gene expression parameters. Input these parameters into the model trained in step 2 to obtain model predictions. Determine the accuracy of the model predictions by comparing the model predictions with the diagnosis results.
[0119] It should be noted that the features and technical effects described in this article for different aspects can be used as reference for each other and will not be repeated here.
[0120] The following examples illustrate this application, but this should not be construed as limiting the scope of the subject matter of this application to the following examples. All technologies implemented based on the above content of this application fall within the scope of this application. The compounds or reagents used in the following examples are commercially available or prepared by conventional methods known to those skilled in the art; the experimental instruments used are commercially available.
[0121] Example 1: Imprinted gene screening and specific probe design
[0122] A total of 79 cervical tissue samples were collected, including 30 benign cervical tissues, 13 cases of CIN1, 14 cases of CIN3, and 22 cases of invasive cervical cancer. All samples were prepared using the conventional formalin-fixed paraffin-embedded (FFPE) method to prepare tissue blocks, cut into 10 μm-thick tissue sections, and adhered to glass slides.
[0123] The eighth intron of the HM13 gene was selected as the target, and a specific probe was designed. The target sequence was NC000020.11:31554829-31559610, strand:+, where strand:+ indicates that the target sequence is a positive strand. Since this embodiment uses the RNAscope method for RNA in situ hybridization, the probe design references Fay Wang, et al. The Journal of Molecular Diagnostics, 2012, 14 (1), 22-29. and U.S. Patent US 7,709,198, and more than three pairs of "double Z" type probes were designed. Each probe contains a sequence of 18-25 nucleotides in length that is complementary to the target sequence, a spacer sequence, and a tail sequence of 14 nucleotides in length. The two tail sequences of each pair of probes together form a sequence of 28 nucleotides in length, which is used for further hybridization to form a branched DNA structure for signal amplification. The optimal temperature for in situ hybridization is typically 15-25°C lower than the Tm value. Setting the Tm value above 60°C allows hybridization detection at 40°C. This hybridization temperature is generally higher than room temperature, avoiding nonspecific hybridization, and significantly reduces concentration changes caused by evaporation of the hybridization solution during prolonged hybridization. Therefore, in this example, the Tm value of the region where the probe and target sequence pair is set to above 60°C. To improve the efficiency of branched DNA formation, according to a scheme similar to RNAscope, with reference to Chinese Patent ZL202110581853.9, the distance between the binding sites of the two probes (first probe and second probe) in the same pair and the target sequence should be less than or equal to 2 nucleotides. To specifically detect nascent RNA expressed by imprinted genes and avoid nonspecific hybridization with other RNA molecules, in this protocol, any first and second probes for the same gene must not simultaneously have consecutive sequences of 15 nucleotides or longer that form continuous complementary pairs with naturally occurring human mRNAs, noncoding RNAs, and introns of other genes. Noncoding RNAs include long noncoding RNA (lncRNA), microRNA (miRNA), ribosomal RNA (rRNA), and transfer RNA (tRNA). Continuous complementary pairs mean that the distance between the binding sites of the two probes on the same nucleotide is less than or equal to 2 nucleotides. Based on these principles, 20 probe pairs for the HM13 gene were designed. The target binding sequences for each probe are shown in Table 1. The linker and tail sequences were designed by Advanced Cell Diagnostics based on U.S. Patent 7,709,198.
[0124] Table 1. HM13 probe target binding sequences
[0125] Following the same principles, probes were designed targeting the imprinted genes SNU13 intron 2 (NC000022.11:41675195-41680243, strand: -), GNAS intron 1 (NC000020.11:58839740-58911195, strand: +), and SNRPN intron 5 (NC000015.10:24962209-24967028, strand: +). Strand: + indicates the target sequence is the positive strand; strand: - indicates the target sequence is the negative strand, which is the reverse complement of the positive strand. The target binding sequences for these three genes are shown in Tables 2-4, respectively.
[0126] Table 2. SNU13 probe target binding sequences
[0127] Table 3. GNAS probe target binding sequences
[0128] Table 4. SNRPN probe target binding sequences
[0129] Using the RNAscope Red Detection Kit, sample pre-treatment, in situ hybridization, and signal amplification were performed according to the detection steps of the kit, and the result images were obtained under an ordinary bright field microscope. Taking the detection results of the HM13 gene allele expression status of an invasive cervical cancer sample as an example, as shown in Figure 2, the cell nucleus appears blue due to hematoxylin staining, and the expression sites of the HM13 gene in the cell nucleus are displayed as red signal points. It can be seen that there are cell nuclei without HM13 gene expression sites, cell nuclei with 1 HM13 gene expression site, cell nuclei with 2 HM13 gene expression sites, and cell nuclei with 3 or more HM13 gene expression sites. The three parameters of biallelic expression (BAE), multiallelic expression (MAE), and total expression (TE) of the imprinted gene HM13 were calculated according to the following formula: Biallelic expression of imprinted gene = N2 / (N1+N2+N 3+ )×100%; Multi-allelic expression of imprinted genes = N 3+ / (N1+N2+N 3+ )×100%; Total expression of imprinted genes = (N1+N2+N 3+ ) / (N0+N1+N2+N 3+ )×100%;
[0130] Among them, N0 is the number of cell nuclei without HM13 gene expression sites, N1 is the number of cell nuclei with one HM13 gene expression site, N2 is the number of cell nuclei with two HM13 gene expression sites, and N 3+ is the number of nuclei with three or more HM13 gene expression sites. The same calculations were performed for the imprinted genes SNU13, GNAS, and SNRPN to obtain the corresponding quantitative parameters. The above formulas were calculated based on counting at least 600 nuclei.
[0131] The results of the allele expression status detection of the HM13 gene are shown in Figure 3. In benign samples, most cells do not express HM13, and a few cells express one HM13 gene. In CIN1 samples, the proportion of cells expressing HM13 increases, and some cells express two HM13 genes. In CIN3 samples, the proportion of cells expressing HM13 further increases, the proportion of cells expressing two HM13 genes increases, and a small number of cells expressing three or more copies of the HM13 gene appear. In invasive cervical cancer samples, the proportions of cells expressing two HM13 genes and cells expressing three or more copies of the HM13 gene are both further increased than those in CIN3.
[0132] Figure 4 shows the results of allelic expression status testing for the SNU13 gene. In benign samples, most cells do not express SNU13, while a small number of cells express one SNU13 gene. In CIN1 samples, the proportion of cells expressing SNU13 increases, and some cells express two SNU13 genes. In CIN3 samples, the proportion of cells expressing SNU13 increases further, as does the proportion of cells expressing two SNU13 genes, with a small number of cells expressing three or more SNU13 gene copies. In invasive cervical cancer samples, the proportions of cells expressing two SNU13 genes and cells expressing three SNU13 gene copies both increase further than in CIN3.
[0133] The results of the allele expression status detection of the GNAS gene are shown in Figure 5. In benign samples, most cells do not express GNAS, and a few cells express one GNAS gene; in CIN1 samples, the proportion of cells expressing GNAS increases, and some cells express two GNAS genes; in CIN3 samples, the proportion of cells expressing GNAS further increases, the proportion of cells expressing two GNAS genes increases, and a small number of cells expressing three or more copies of the GNAS gene appear; in invasive cervical cancer samples, the proportions of cells expressing two GNAS genes and cells expressing three or more copies of the GNAS gene are both further increased than those in CIN3.
[0134] The results of the allele expression status detection of the SNRPN gene are shown in Figure 6. In benign samples, most cells do not express SNRPN, and a few cells express one SNRPN gene; the proportion of cells expressing SNRPN in CIN1 samples increased significantly, and some cells expressed two SNRPN genes, and some cells even expressed three or more copies of the GNAS gene; in CIN3 samples, the proportion of cells expressing two SNRPN genes and cells expressing three or more copies of the SNRPN gene increased; in invasive cervical cancer samples, the proportion of cells expressing two SNRPN genes and cells expressing three or more copies of the SNRPN gene increased further than that in CIN3.
[0135] Because benign and CIN1 cases do not require specific clinical treatment, while CIN3 and invasive cervical cancer cases require cervical conization or other surgical intervention, the 30 benign cervical tissue samples and 13 CIN1 samples were combined into a low-risk group, and the 14 CIN3 and 22 invasive cervical cancer samples were combined into a high-risk group. Statistical analysis of the BAE, MAE, and TE parameters of the low- and high-risk groups was performed. As shown in Figure 7, the BAE, MAE, and TE parameters of the GNAS, HM13, and SNU13 genes in the high-risk group were all significantly increased compared to the low-risk group. Furthermore, for the SNRPN gene, only the MAE was significantly increased in the high-risk group compared to the low-risk group, while the BAE and TE parameters were not significantly increased.
[0136] ROC curves were plotted for the BAE, MAE, and TE parameters of the GNAS, SNRPN, HM13, and SNU13 genes. As shown in Figure 8, the areas under the ROC curves for the BAE parameter of the GNAS, HM13, and SNU13 genes were all greater than 0.8, while the area under the ROC curve for the BAE parameter of the SNRPN gene was only 0.61. The areas under the ROC curves for the MAE parameter of the GNAS, HM13, and SNU13 genes were all greater than 0.8, while the area under the ROC curve for the BAE parameter of the SNRPN gene was only 0.71. The areas under the ROC curves for the TE parameter of the GNAS, HM13, and SNU13 genes were all greater than 0.75, while the area under the ROC curve for the BAE parameter of the SNRPN gene was only 0.59. Statistical analysis also showed that the areas under the ROC curves for the BAE, MAE, and TE parameters of the GNAS, HM13, and SNU13 genes were significantly different from those of the corresponding parameters of the SNRPN gene. Therefore, GNAS, HM13 and SNU13 genes are significantly better than SNRPN gene in distinguishing benign cervical, CIN1 samples from CIN3, invasive cervical cancer samples.
[0137] Example 2: Establishment of a cervical cancer risk prediction model
[0138] According to the gene screening results in Example 1, HM13, SNU13 and GNAS genes were selected for training the cervical cancer risk prediction model in cervical cell scraping samples.
[0139] 75 patients with high-risk HPV infection were selected, including 29 benign cases, 15 CIN1 cases, 15 CIN3 cases, and 16 invasive cervical cancer cases. Cervical cell samples were obtained from 75 patients using cervical cell scraping. The cell samples were placed in 10% neutral formalin and fixed at room temperature for 24-48 hours. The cells were collected by centrifugation and adhered to a glass slide. In situ hybridization detection was performed using the RNAscope multi-channel fluorescence detection kit, and the target binding sequence of the probe used was the same as in Example 1. After the experiment, the result image was collected using a fluorescence microscope. As shown in Figure 9, cell nuclei containing different numbers of signal points can be observed.
[0140] The biallelic expression (BAE) and tri- to quadri-allelic expression (MAE) of the imprinted gene HM13 were calculated according to the following formulas: 3-4 ), penta-hexa-allelic expression (MAE 5-6 ), seven to eight allele expression (MAE 7-8 ), more than nine alleles expressed (MAE 9+ ) and total expression (TE): Biallelic expression of imprinted genes = N2 / (N1+N2+N3+N4+N5+N6+N7+N8+N 9+ )×100%; Expression of three to four alleles of imprinted genes = (N3+N4) / (N1+N2+N3+N4+N5+N6+N7+N8+N 9+ )×100%; Expression of five to six alleles of imprinted genes = (N5+N6) / (N1+N2+N3+N4+N5+N6+N7+N8+N 9+ )×100%; Expression of imprinted gene alleles seven to eight = (N7+N8) / (N1+N2+N3+N4+N5+N6+N7+N8+N 9+ )×100%; Expression of more than nine alleles of imprinted genes = N 9+ / (N1+N2+N3+N4+N5+N6+N7+N8+N 9+ )×100%; Total expression of imprinted genes = (N1+N2+N3+N4+N5+N6+N7+N8+N 9+ ) / (N0+N1+N2+N3+N4+N5+N6+N7+N8+N 9+ ) × 100%;
[0141] Among them, N0 is the number of cell nuclei without HM13 gene expression sites, N1 is the number of cell nuclei with 1 HM13 gene expression site, N2 is the number of cell nuclei with 2 HM13 gene expression sites, N3 is the number of cell nuclei with 3 HM13 gene expression sites, N4 is the number of cell nuclei with 4 HM13 gene expression sites, N5 is the number of cell nuclei with 5 HM13 gene expression sites, N6 is the number of cell nuclei with 6 HM13 gene expression sites, N7 is the number of cell nuclei with 7 HM13 gene expression sites, N8 is the number of cell nuclei with 8 HM13 gene expression sites, and N 9+ is the number of cell nuclei with 9 or more HM13 gene expression sites. The above calculations were also performed for the imprinted genes SNU13 and GNAS genes to obtain the corresponding quantitative parameters. The above formulas were calculated after counting at least 200 cell nuclei. According to the above formulas, TE is still the ratio of all cell nuclei with imprinted gene expression signals to all cell nuclei, BAE is still the ratio of all cell nuclei with two imprinted gene expression signals to all cell nuclei with imprinted gene expression signals, and the calculation method is the same as in Example 1, while MAE 3-4 、MAE 5-6 、MAE 7-8 、MAE 9+ In essence, the MAE in Example 1 was further decomposed to more accurately quantify the abnormal expression of imprinted genes.
[0142] The results of the allele expression status detection of the GNAS gene are shown in Figure 10. In benign samples, most cells do not express GNAS, and a few cells express one GNAS gene; the proportion of cells expressing GNAS in CIN1 samples increases, and some cells express two GNAS genes; the proportion of cells expressing GNAS in CIN3 samples further increases, the proportion of cells expressing two GNAS genes increases, and a small number of cells expressing three or more copies of the GNAS gene appear; in invasive cervical cancer samples, the proportions of cells expressing two GNAS genes and cells expressing three or more copies of the GNAS gene are both further increased than those in CIN3.
[0143] The results of the allele expression status detection of the HM13 gene are shown in Figure 11. In benign samples, most cells do not express HM13, and a few cells express one HM13 gene; the proportion of cells expressing HM13 in CIN1 samples increases, and some cells express two HM13 genes; the proportion of cells expressing HM13 in CIN3 samples further increases, the proportion of cells expressing two HM13 genes increases, and a small number of cells expressing three or more copies of the HM13 gene appear; in invasive cervical cancer samples, the proportions of cells expressing two HM13 genes and cells expressing three or more copies of the HM13 gene are further increased than those in CIN3.
[0144] Figure 12 shows the results of allelic expression status testing for the SNU13 gene. In benign samples, most cells do not express SNU13, while a small number of cells express one SNU13 gene. In CIN1 samples, the proportion of cells expressing SNU13 increases, and some cells express two SNU13 genes. In CIN3 samples, the proportion of cells expressing SNU13 increases further, as does the proportion of cells expressing two SNU13 genes, with a small number of cells expressing three or more SNU13 gene copies. In invasive cervical cancer samples, the proportions of cells expressing two SNU13 genes and cells expressing three SNU13 gene copies both increase further than in CIN3.
[0145] Based on the results of colposcopic biopsies of 75 samples, the receiver operating characteristic (ROC) curves of the six quantitative parameters of the three genes HM13, SNU13, and GNAS were drawn and the area under the curve (AUROC) was calculated. The AUROC results are shown in Table 5.
[0146] Table 5. Area under the ROC curve (AUROC) for quantitative parameters of imprinted genes HM13, SNU13, and GNAS in the training set
[0147] 1. Establishment of decision tree model
[0148] Single-gene decision tree models were constructed for HM13, SNU13, and GNAS using the rpart function in the R language rpart package. Decision tree models were also constructed for combinations of at least two genes from HM13, SNU13, and GNAS. The minsplit parameter was set to 10, meaning that only nodes with at least 10 samples were subject to decision criteria and the next step was not considered. Nodes with fewer than 10 samples were not subject to the next step. The maxdepth parameter was set to 3, meaning that a maximum of three layers of decision criteria were set. The quantitative parameters for imprinted genes used in each model were automatically selected by the program. The selected quantitative parameters and the area under the receiver operating characteristic (ROC) curve for each model are shown in Table 6.
[0149] Table 6. Quantitative parameters of the decision tree model and the area under the ROC curve in the training set
[0150] The structure of the decision tree and the threshold value of the conditional judgment are automatically selected by the program. The judgment conditions of each decision tree model are as follows:
[0151] When the GNAS gene is used alone, the condition for judging a sample as negative is that the TE of the GNAS gene is less than 21%, and the condition for judging a sample as positive is that the TE of the GNAS gene is greater than or equal to 21%.
[0152] When the HM13 gene is used alone, the condition for judging the sample as negative is the MAE of the HM13 gene 7-8 Less than 0.28% and MAE of HM13 gene 3-4 Less than 7.7%, or MAE of the HM13 gene 7-8 Less than 0.28% MAE of HM13 gene 3-4 The condition for judging the sample as positive is that the MAE of the HM13 gene is greater than or equal to 7.7% and the BAE of the HM13 gene is less than 19%. 7-8 Less than 0.28% MAE of HM13 gene 3-4 Greater than or equal to 7.7% and BAE of HM13 gene greater than or equal to 19%, or MAE of HM13 gene 7-8 Greater than or equal to 0.28%.
[0153] When the SNU13 gene is used alone, the condition for judging the sample as negative is the MAE of the SNU13 gene. 3-4 Less than 5.3%, or MAE of the SNU13 gene 3-4 The condition for judging a sample as positive is that the MAE of the SNU13 gene is greater than or equal to 5.3%, the BAE of the SNU13 gene is less than 23.5%, and the TE of the SNU13 gene is greater than or equal to 15.4%. 3-4 greater than or equal to 5.3%, the BAE of the SNU13 gene is less than 23.5% and the TE of the SNU13 gene is less than 15.4%, or the MAE of the SNU13 gene 3-4 greater than or equal to 5.3% and the BAE of the SNU13 gene was greater than or equal to 23.5%.
[0154] When using the combination of GNAS and HM13 genes, the conditions for judging the sample as negative are that the TE of the GNAS gene is less than 21% and the MAE of the HM13 gene is less than 21%. 3-4 Less than 7.7%, or TE of GNAS gene less than 21%, MAE of HM13 gene 3-4Greater than or equal to 7.7% and BAE of HM13 gene less than 26%, or TE of GNAS gene greater than or equal to 21% and MAE of HM13 gene 5-6 The conditions for judging a sample as positive are that the TE of the GNAS gene is less than 21% and the MAE of the HM13 gene is less than 0.095%. 3-4 Greater than or equal to 7.7% and BAE of HM13 gene greater than or equal to 26%, or TE of GNAS gene greater than or equal to 21% and MAE of HM13 gene 5-6 Greater than or equal to 0.095%.
[0155] When using the combination of GNAS and SNU13 genes, the conditions for judging the sample as negative are that the TE of the GNAS gene is less than 21% and the MAE of the SNU13 gene is less than 21%. 3-4 Less than 5.3%, or TE of GNAS gene less than 21%, MAE of SNU13 gene 3-4 The conditions for judging the sample as positive are that the TE of GNAS gene is less than 21% and the MAE of SNU13 gene is greater than or equal to 5.3% and the TE of SNU13 gene is greater than or equal to 15.4%. 3-4 greater than or equal to 5.3% and the TE of the SNU13 gene is less than 15.4%, or the TE of the GNAS gene is greater than or equal to 21%.
[0156] When using a combination of HM13 and SNU13 genes, the condition for judging a sample as negative is the MAE of the SNU13 gene. 3-4 Less than 5.3%, or MAE of the SNU13 gene 3-4 MAE greater than or equal to 5.3% and HM13 3-4 Less than 7.7%, the condition for judging the sample as positive is the MAE of SNU13 gene 3-4 MAE greater than or equal to 5.3% and HM13 3-4 Greater than or equal to 7.7%.
[0157] When using the combination of GNAS, HM13, and SNU13 genes, the condition for judging the sample as negative is that the TE of the GNAS gene is less than 21% and the MAE of the SNU13 gene is less than 21%. 3-4 Less than 5.3%, or TE of GNAS gene less than 21%, MAE of SNU13 gene 3-4 Greater than or equal to 5.3% and TE of HM13 gene less than 23%, or TE of GNAS gene greater than or equal to 21% and MAE of HM13 gene 5-6 The conditions for judging a sample as positive are that the TE of the GNAS gene is less than 21% and the MAE of the SNU13 gene is less than 0.095%. 3-4Greater than or equal to 5.3% and TE of HM13 gene greater than or equal to 23%, or TE of GNAS gene greater than or equal to 21% and MAE of HM13 gene 5-6 Greater than or equal to 0.095%.
[0158] According to the above model structure, the diagnostic accuracy of each model in the training set is shown in Table 7.
[0159] Table 7. Sensitivity and specificity of the decision tree model in the training set
[0160] 2. Establishment of Logistic Regression Model
[0161] Logistic regression models were established for HM13, SNU13, and GNAS genes using the glm function of the R language stats package. Each gene was analyzed from the BAE, MAE, 3-4 、MAE 5-6 、MAE 7-8 、MAE 9+ Four of the six quantitative parameters, including TE, were selected to establish the model. The parameters selected for each gene were automatically selected by the rfe function and rfeControl function of the caret package in R language according to the discrimination after parameter combination. After automatic selection, the quantitative parameters used in the GNAS gene model were BAE, MAE 3-4 、MAE 5-6 and TE, the quantitative parameter used in the HM13 gene model is MAE 3-4 、MAE 5-6 、MAE 7-8 The quantitative parameters used for the TE gene model are SNU13: BAE, MAE 3-4 、MAE 5-6 and TE.
[0162] The logistic regression curve function of the GNAS gene is:
[0163] The logistic regression curve function of the HM13 gene is:
[0164] The logistic regression curve function of the SNU13 gene is:
[0165] Where Probability represents the probability that the sample is CIN3 or invasive cervical cancer, Exp represents the exponential function of the base of the natural logarithm, e, and the exponent of e is in the brackets after Exp.
[0166] Based on the above models, the areas under the receiver operating characteristic (ROC) curves for the single-gene models for GNAS, HM13, and SNU13 in the training set are shown in Table 8. If the probability threshold is set at 85%, with a minimum overall sensitivity for CIN3 and invasive cervical cancer, with a probability greater than or equal to the threshold considered positive and a probability less than the threshold considered negative, the sensitivity and specificity of the single-gene models for GNAS, HM13, and SNU13 in the training set are shown in Table 8.
[0167] Table 8. Area under the ROC curve, probability threshold, sensitivity and specificity of the single-gene logistic regression model in the training set
[0168] When combining the prediction results of two or three genes, the predicted probabilities of each gene are multiplied by different coefficients and then added together. To select the optimal combination coefficient, the coefficient is increased or decreased by 0.1, and the combination with the highest accuracy is selected through exhaustive selection.
[0169] The results of combining the two genes GNAS and HM13 with different coefficients are shown in Table 9:
[0170] Table 9. Area under the ROC curve, probability threshold, sensitivity and specificity of the two genes GNAS and HM13 in the training set after different coefficient combinations
[0171] Table 9 shows that all combinations achieved a sensitivity of 87.1%. Specificity reached its highest level, reaching 72.7%, when the GNAS coefficient was 0.6, the HM13 coefficient was 0.4, and when both the GNAS and HM13 coefficients were 0.5. Given the same sensitivity and specificity, the area under the receiver operating characteristic curve (ROC) curve for a GNAS and HM13 coefficient of 0.5 was greater than that for a GNAS coefficient of 0.6 and an HM13 coefficient of 0.4. Therefore, a GNAS and HM13 coefficient of 0.5 was selected as the optimal combination, with a probability threshold of 0.461.
[0172] The results of combining the two genes GNAS and SNU13 with different coefficients are shown in Table 10:
[0173] Table 10. Area under the ROC curve, probability threshold, sensitivity and specificity of the two genes GNAS and SNU13 when combined with different coefficients in the training set
[0174] As shown in Table 10, the sensitivity of all combinations is 87.1%. When the coefficient of GNAS is 0.1 and the coefficient of SNU13 is 0.9, the specificity is the highest, which is 65.9%. Therefore, the coefficient of GNAS is 0.1 and the coefficient of SNU13 is 0.9 as the optimal combination condition, and the probability threshold is 0.511.
[0175] The results of combining the two genes HM13 and SNU13 with different coefficients are shown in Table 11:
[0176] Table 11. Area under the ROC curve, probability threshold, sensitivity and specificity of the two genes HM13 and SNU13 when combined with different coefficients in the training set
[0177] As shown in Table 11, the sensitivity of all combinations is 87.1%. When the coefficient of HM13 is 0.6 and the coefficient of SNU13 is 0.4, the specificity is the highest, which is 65.9%. Therefore, the coefficient of HM13 is 0.6 and the coefficient of SNU13 is 0.4 as the optimal combination condition, and the probability threshold is 0.418.
[0178] The results of different combinations of the three genes GNAS, HM13 and SNU13 are shown in Table 12:
[0179] Table 12. Area under the ROC curve, probability threshold, sensitivity and specificity of the three genes GNAS, HM13 and SNU13 in the training set after different coefficient combinations
[0180] Table 12 shows that all combinations achieved a sensitivity of 87.1%. Specificity reached its highest level of 75.0% when the GNAS coefficient was 0.4, the HM13 coefficient was 0.4, and the SNU13 coefficient was 0.2, and when the GNAS coefficient was 0.5, the HM13 coefficient was 0.4, and the SNU13 coefficient was 0.1. Given the same sensitivity and specificity, the area under the receiver operating characteristic curve for the GNAS coefficients of 0.4, the HM13 coefficients of 0.4, and the SNU13 coefficients of 0.2 was greater than that for the GNAS coefficients of 0.5, the HM13 coefficients of 0.4, and the SNU13 coefficients of 0.1. Therefore, the optimal combination of GNAS coefficients of 0.4, HM13 coefficients of 0.4, and SNU13 coefficients was selected, with a probability threshold of 0.466.
[0181] Example 3: Evaluation of cervical cancer risk prediction model
[0182] 105 patients with high-risk HPV infection were enrolled consecutively. Cervical cell samples were obtained from 105 patients by cervical cell scraping. The sample processing method and RNAscope detection method were the same as in Example 2. The biallelic expression (BAE) and tri- to tetra-allelic expression (MAE) of the imprinted genes GNAS, HM13, and SNU13 were calculated according to the formula in Example 2. 3-4 ), penta-hexa-allelic expression (MAE 5-6 ), seven to eight allele expression (MAE 7-8 ), more than nine alleles expressed (MAE 9+ ) and total expression (TE) parameters.
[0183] The personnel involved in sampling, experiments, and model prediction were all unaware of the pathological results of the samples. The pathological results of subsequent colposcopic biopsies showed that the 105 samples included in this example included 49 benign cases, 24 CIN1 cases, 16 CIN3 cases, and 16 invasive cervical cancers.
[0184] 1. Verification of the decision tree model
[0185] The validation set samples were judged as negative or positive based on the criteria of the decision tree model in Example 2. The validation results for the models of individual genes, GNAS, HM13, and SNU13, and for any two or all three of these genes are shown in Table 13.
[0186] Table 13. Sensitivity and specificity of the decision tree model in the validation sample
[0187] The validation results showed that when the HM13 gene was used alone, the sensitivity was 90.6% and the specificity was 72.6%; when HM13 and SNU13 were used in combination, the sensitivity was 78.1% and the specificity was 93.2%; when the combination of GNAS, HM13 and SNU13 was used, the sensitivity was 81.3% and the specificity was 84.9%.
[0188] 2. Validation of the Logistic Regression Model
[0189] The corresponding parameters of the imprinted genes GNAS, HM13, and SNU13 from the validation set samples were substituted into the single-gene logistic regression function to calculate the probabilities. The probabilities of GNAS, HM13, and SNU13 were directly used, or any two or three of these genes were combined according to the optimal combination coefficient. Predictions were made for the validation samples using the probability thresholds described in Example 2. The predicted results were compared with the pathological results of subsequent colposcopic biopsies of the validation samples. The results are shown in Table 14:
[0190] Table 14. Sensitivity and specificity of the logistic regression model in the validation sample
[0191] The validation results showed that when the HM13 gene was used alone, the sensitivity was 93.8% and the specificity was 67.1%; when GNAS was added to HM13, the sensitivity was 90.6% and the specificity increased to 76.7%; when the SNU13 gene was used alone, the sensitivity was 87.5% and the specificity was 83.6%; when GNAS was added to SNU13, the sensitivity was 87.5% and the specificity was 82.2%; when HM13 and SNU13 were used in combination, the sensitivity reached 96.9% and the specificity was 74.0%; when the three genes HM13, SNU13 and GNAS were used simultaneously, the sensitivity reached 93.8% and the specificity reached 83.6%.
[0192] The above validation results demonstrate that in situ hybridization using intronic probes to detect the expression of the imprinted genes HM13, SNU13, and GNAS can effectively differentiate between low-risk benign lesions and CIN1 and high-risk CIN3 and invasive cervical cancer. Both decision tree and logistic regression models achieved high sensitivity and specificity. The high specificity of the detection method of the present invention suggests that, for patients positive for high-risk HPV infection, detecting the expression of the imprinted genes HM13, SNU13, and GNAS can avoid unnecessary colposcopy and biopsy in 67.1% to 93.2% of benign and CIN1 cases, effectively reducing patient suffering and conserving significant medical resources. The high sensitivity, particularly for CIN3 samples, reaches 75% to 93.8%, a significant improvement over other detection technologies with sensitivities of less than 70%. This can effectively prevent missed diagnoses of CIN3 samples, prompting patients to undergo timely treatment such as cervical conization to prevent progression to invasive cervical cancer, which carries a worse prognosis.
[0193] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0194] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A gene marker, characterized in that, Comprising at least one of the following genes: HM13 and SNU13; Optionally, the gene marker further comprises the GNAS gene.
2. Use of a reagent for detecting the gene marker according to claim 1 in the preparation of a kit, the reagent being used for diagnosing cervical intraepithelial neoplasia and invasive cervical cancer.
3. The use according to claim 2, wherein The reagent comprises at least one of a probe, a primer, and a mass spectrometry detection reagent that are specific for the gene marker.
4. The use according to claim 3, characterized in that, The probe has a nucleotide sequence shown in SEQ ID NO: 1 to 144.
5. A kit for screening biological samples for cervical intraepithelial neoplasia and invasive cervical cancer, characterized in that, Comprising: A reagent suitable for detecting at least one of the gene markers according to claim 1.
6. The kit according to claim 5, characterized in that, The reagent comprises at least one of a probe, a primer, and a mass spectrometry detection reagent that are specific for the gene marker; Optionally, the reagent comprises a probe specific for the gene marker, and the probe has a nucleotide sequence shown in SEQ ID NO: 1 to 144; Optionally, the kit further comprises at least one of a fluorescent reagent, a buffer, and a washing solution.
7. A drug for treating cervical intraepithelial neoplasia and invasive cervical cancer, characterized in that, The drug contains: A reagent that specifically alters the gene marker according to claim 1.
8. The drug according to claim 7, wherein The reagent is a reagent based on a gene editing or nucleic acid synthesis method; Optionally, the drug further comprises: a pharmaceutically acceptable carrier or excipient.
9. An intron probe, characterized in that, The probe is used for detecting the gene marker according to claim 1.
10. The probe according to claim 9, wherein The probe has a nucleotide sequence shown in SEQ ID NO: 1 to 144.
11. A method for training a machine learning model, the machine learning model being used to predict the risk of cervical intraepithelial neoplasia and invasive cervical cancer, characterized in that, Including: Obtaining intron in situ hybridization image information of a gene marker in a training sample, the image information containing hybridization signals, and the hybridization signals are obtained by hybridizing the probe according to claim 9 or 10 with the gene marker according to claim 1; Based on the image information, obtaining a predetermined feature of the training sample; Inputting the predetermined feature into a machine learning model, using the known risk status of the training sample as a label to perform supervised training on the machine learning model to obtain the machine learning model.
12. The method according to claim 11, wherein The machine learning model is selected from at least one of logistic regression, neural network, decision tree, and random forest; Preferably, the machine learning model is selected from logistic regression.
13. The method according to claim 11, wherein The predetermined feature includes at least one selected from the proportion of cells with two hybridization signals among all cells, the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, the proportion of cells with seven and eight hybridization signals among all cells, the proportion of cells with more than nine hybridization signals among all cells, and the proportion of cells with hybridization signals among all cells.
14. The method according to claim 13, wherein The predetermined feature includes at least one selected from the proportion of cells with two hybridization signals among all cells, the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, the proportion of cells with seven and eight hybridization signals among all cells, and the proportion of cells with hybridization signals among all cells.
15. The method according to claim 11 or 14, characterized in that, The predetermined characteristics of the GNAS gene and the SNU13 gene are selected from at least one of the proportion of cells with two hybridization signals among all cells, the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, and the proportion of cells with hybridization signals among all cells.
16. The method according to claim 11 or 14, characterized in that, The predetermined characteristics of the HM13 gene are selected from at least one of the proportion of cells with three and four hybridization signals among all cells, the proportion of cells with five and six hybridization signals among all cells, the proportion of cells with seven and eight hybridization signals among all cells, and the proportion of cells with hybridization signals among all cells.
17. The method according to claim 11, wherein The image information is obtained by a quantitative chromogenic blotting gene in situ hybridization method.
18. A risk prediction model for cervical cancer occurrence, characterized in that, Including: An image information acquisition module for acquiring the image information of the gene marker of the sample to be tested as described in claim 1; A predetermined characteristic acquisition module for acquiring the predetermined characteristics of the training sample based on the image information; And A prediction module for inputting the predetermined characteristics into a pre-trained machine learning model to obtain a prediction result; wherein the pre-trained machine learning model is trained by the method described in any one of claims 9 to 15; 19. The prediction model according to claim 18, wherein The sample to be tested is selected from cervical cell tissues.
20. The prediction model according to claim 18, wherein The determination method of the prediction result is as follows: If the output result of the machine learning model is greater than or equal to 0.466, the sample is determined to be a positive sample of cervical cancer; If the output result of the machine learning model is less than 0.466, the sample is determined to be a positive sample of cervical cancer.
21. A server, characterized in that, The server includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the machine learning model training method described in any one of claims 11 to 17 is implemented.
22. A computer-readable storage medium containing a computer program, characterized in that, When the computer program is executed by one or more processors, the machine learning model training method described in any one of claims 11 to 17 is implemented.
23. A method for predicting or diagnosing the risk of cervical cancer, characterized in that, Including: Obtaining the intron in situ hybridization image information of the gene marker of the sample to be tested, the image information containing hybridization signals, which are obtained by hybridizing the probe described in claim 9 or 10 with the gene marker described in claim 1; Based on the image information, obtaining the predetermined characteristics of the sample to be tested; Inputting the predetermined characteristics into a pre-trained machine learning model to obtain a prediction result; the pre-trained machine learning model is trained by the method described in any one of claims 11 - 17.
24. The method according to claim 23, wherein The predetermined characteristics input into the pre-trained machine learning model may further include text characteristics; Optionally, the text characteristics include at least one of the age information, body index, and past medical history of the sample to be tested.
25. The method according to claim 23, wherein The sample to be tested is selected from cervical cell tissues.
26. The method according to claim 23, wherein The determination method of the prediction result is as follows: If the output result of the machine learning model is greater than or equal to a predetermined threshold, the sample is determined to be a positive sample of cervical cancer; if the output result of the machine learning model is less than the predetermined threshold, the sample is determined to be a positive sample of cervical cancer; Optionally, the predetermined threshold is 0.4 - 0.5, preferably 0.
466.
27. The method according to claim 26, wherein After determining whether the sample is negative or positive through the above determination method, further examinations and treatments including medications and surgeries can be performed on cervical lesions; Optionally, negative cases can undergo liquid-based cytology examinations every three months or interferon antiviral therapy; Optionally, for positive cases, further colposcopy and colposcopic biopsy are required. If it is determined to be CIN3, ablation therapies such as laser, electrocautery, and cryotherapy can be performed, or cervical conization can be carried out; Optionally, if it is determined to be invasive cervical cancer, cervical conization or radical hysterectomy for cervical cancer is required, and radiotherapy or chemotherapy may be performed as needed to eliminate metastatic cancer cells.
Citation Information
Patent Citations
Method and agents to quantify proteins from tissues
CA2875418A1
Grading model for detecting benign and malignant degrees of tumors and application thereof
CN116130099A
Therapeutic oligonucleotides
US20190256546A1
Modulation of novel immune checkpoint targets
US20200016202A1
Determining risk of cancer recurrence
US20230026291A1